Repair: Null-Out
On this page
Summary
Null-out is the simplest and most honest repair: samples flagged as Bad hole, plus a few samples of padding on each side, are set to null and no value is invented. It is the default. The cost is gaps in the curve, which every later calculation has to deal with.
Inputs and outputs
| Item | Units | |
|---|---|---|
| Input | A curve to be repaired, for example Bulk density | curve units |
| Input | A Bad-hole flag curve and a padding length | 0 or 1; ft |
| Output | The curve with flagged and padded samples set to null | curve units |
Equations
The unpadded flag \(\badFlag\) is grown by \(m\) samples on each side. With \(\Delta z\) the sample step and \(L_{p}\) the padding length:
This is a morphological dilation of the flag. The repaired curve is the original with the padded samples removed:
The fraction of the interval lost is the share of padded samples, and it rises with the padding length and with the number of separate flagged intervals. A flag that is a few isolated samples loses \(2m+1\) samples for each.
| Symbol | Variable | Units | Typical range |
|---|---|---|---|
| \(F_{bh}\) | Bad-hole flag | ||
| \(\rho_b\) | Bulk density | g/cm³ | 1.8 to 3.0 |
| Bad hole |
Single-value calculator
No calculator: nulling is a procedure on a depth series. The worked example pads a flag and measures what remains.
Behavior
The caliper flags only the part of a washout where the hole is already far over gauge. The density pad starts to lose contact before that point and recovers after it, so the unflagged samples next to the flag are still damaged. In the worked example, with a 100 ft interval at 0.5 ft sampling and one 12 ft washout of which the caliper flags 9.5 ft (19 samples), the worst remaining density error falls from 0.272 g/cm³ with no padding to 0.194 with a padding of 2 samples, 0.117 with 4, 0.039 with 6 and none at 10. The cost is the data lost: 19 samples with no padding, 27 with 4 (13.5% of the interval) and 39 with 10. The zone mean density is 2.5546 g/cm³ for the true curve, 2.5126 for the raw curve and 2.5517 after nulling with a padding of 4 samples (2 ft): the damaged samples bias a zone average by 0.04 g/cm³ and nulling removes nearly all of it.
Parameter guidance
Padding length. Pad by at least the vertical resolution of the tool and the distance over which contact is lost and recovered: typically 1 to 5 ft on each side for the density pad. Start with 2 ft and look at the curves beside the flag. A longer padding removes more of the damaged edge but also more good data. The padding can also be set as a factor of the flag length for long washouts.
What to null. Null the curves that the hole damages (density, PE, neutron if the hole is bad, often the sonic is less affected) and leave the others. Null the dependent curves too: a density porosity or a mineral volume calculated from a bad density is as bad as the density. The flags for each curve are in Caliper-Based Flags and Log-Quality Flags, and exempt coal and salt as on Coal and Salt Identification.
After nulling. Decide for each later calculation what to do with a null: skip it, carry it forward and leave a gap in the result, or fill it. If a gap in the output is unacceptable, use a repair from Repair: Regression-Based or Repair: Machine-Learning Infill, and keep the flag so that the repaired intervals can be identified later.
Keep the original. Write the nulled curve to a new curve and keep the raw one. Nulling is irreversible if the raw curve is overwritten.
Worked example
A synthetic 100 ft interval with a smooth density that is damaged around a caliper washout (a smoothed error of up to 0.35 g/cm³, wider than the caliper flag). The flag is padded by 0 to 10 samples and the damage that remains and the data lost are measured. The true density is known because the data are synthetic.
rng = np.random.default_rng(6)
step = 0.5
depth = np.arange(0, 100, step)
# True bulk density and a caliper with one washout
true_rhob = 2.55 + 0.05 * np.sin(depth / 7.0) + rng.normal(0, 0.01, depth.size)
cali = 8.5 + rng.normal(0, 0.03, depth.size)
w0, w1 = 40, 52
cali[(depth >= w0) & (depth < w1)] += 2.5 * np.sin(np.pi * (depth[(depth >= w0) & (depth < w1)] - w0) / (w1 - w0))
# The density pad loses contact earlier than the caliper exceeds the threshold:
# the error is smoothed over about 4 ft, not a step.
err_core = np.where((depth >= w0) & (depth < w1), -0.35, 0.0)
kernel = np.ones(9) / 9
err = np.convolve(err_core, kernel, mode="same")
rhob = true_rhob + err
flag = (cali - 8.5) > 0.75
print(f"flagged: {flag.sum()} samples ({depth[flag].min():.1f}-{depth[flag].max():.1f} ft)")
def pad(f, n):
"""Grow a flag by n samples on each side."""
return np.convolve(f.astype(float), np.ones(2 * n + 1), mode="same") > 0
print(f"\n{'pad (samples)':>13} {'nulled':>7} {'worst remaining error (g/cm3)':>30} {'mean |error| of kept':>22}")
for n in (0, 2, 4, 6, 10):
f = pad(flag, n)
kept_err = np.abs(err[~f])
print(f"{n:13d} {f.sum():7d} {kept_err.max():30.3f} {kept_err.mean():22.4f}")
f = pad(flag, 4)
nulled = np.where(f, np.nan, rhob)
print(f"\nzone mean RHOB: true {true_rhob.mean():.4f} raw {rhob.mean():.4f} null-out (pad 4) {np.nanmean(nulled):.4f}")
print(f"data lost with pad 4: {f.mean():.1%} of the interval")
Output
flagged: 19 samples (41.5-50.5 ft)
pad (samples) nulled worst remaining error (g/cm3) mean |error| of kept
0 19 0.272 0.0105
2 23 0.194 0.0055
4 27 0.117 0.0020
6 31 0.039 0.0002
10 39 0.000 0.0000
zone mean RHOB: true 2.5546 raw 2.5126 null-out (pad 4) 2.5517
data lost with pad 4: 13.5% of the interval
Assumptions and limitations
- The flag marks the whole damaged interval after padding. Damage beyond the padding stays in the curve.
- Padding is short enough to leave useful data between nearby flagged intervals. Two flags separated by less than twice the padding merge into one.
- Later steps tolerate gaps. Many calculations propagate nulls and produce a gap in the result, and some silently skip them, which biases zone averages.
- A bad sample is worse than no sample. This is true for most analyses but not for every one: a rough value with a known uncertainty is sometimes better than nothing.
QC checks
- Nulled intervals are plotted over the original: gaps cover the whole damaged zone and not the good data around it.
- The share of data lost is recorded per well and per formation. A large share means the well needs a repair or a decision that the log is unusable.
- The statistics of the curve in a clean zone before and after nulling are the same, and in a damaged zone the low tail has gone.
- Zone averages and cross-plots no longer show the low-density tail of the washouts.
- The raw and nulled curves are both saved, and the flag, the padding and the thresholds used are recorded.
Going Deeper
Removal of bad data is a much older practice than repair, and it avoids the main problem of any repair: the repaired values look like measurements. Data that is removed leaves visible evidence of the gap, and repaired data does not unless it is flagged. For this reason the default in a careful workflow is to remove, and to repair only where a downstream result needs a continuous curve. When a repair is needed, it is applied to the nulled curve so that the unpadded damaged edges are never used for training or for a taper. The padding is the only subjective parameter, and it is best chosen by looking at what happens at a few washouts in the project and not by a fixed rule.
References
- Asquith, G. and Krygowski, D., 2004. Basic Well Log Analysis, 2nd edition. AAPG Methods in Exploration Series 16, American Association of Petroleum Geologists, Tulsa, OK.
Python reference implementation
Python reference implementation
The Python reference implementation is available to registered users with a verified email address. Register or sign in to view it.