CamPetro

Null and Invalid Value Handling

On this page

Summary

A Null sentinel value such as -999.25 is a flag and not a measurement, and one left in a curve produces large, silent errors in every average, cross-plot and calculation. This page describes how to find the sentinels in a file, convert them to missing values, tell structural gaps from interior ones, and decide what may be filled. Use it on every curve before any statistic is computed.

Inputs and outputs

Item Notes
Input Curve values As read from the file
Input The null value from the file header Usually present, sometimes wrong or different from the data
Input A list of conventional sentinel values For when the header is silent
Input Plausible range of the curve family Optional, for out-of-range values
Output The curve with true missing values NaN, not a number in the data
Output A gap report Leading, trailing and interior gaps with their length
Output Optionally, a gap-filled curve Short interior gaps only

Equations

There is no governing equation, apart from the interpolation used to fill short gaps. The procedure is:

  1. Read the null value from the file header, and use it first. Do not assume that it is correct.
  2. Look in the data for exact repeats of the conventional sentinels (-999.25, -999, -9999, -99999, 9999, and very large numbers such as \(\pm 10^{30}\)), allowing for the file's rounding. Require a minimum number of repeats for the conventional values, because one coincidental sample is not a convention.
  3. Convert every sentinel to a true missing value (NaN). Never compute with the sentinel.
  4. Check ranges. A value outside the physical range of the family (negative resistivity, density below the density of the fluid, gamma ray above the tool limit) is invalid even if it is not a sentinel, and is treated as missing or flagged, depending on the policy.
  5. Classify each run of missing samples as leading (before the first valid sample), trailing (after the last) or interior.
  6. Decide what to fill. Leading and trailing gaps are outside the logged interval and are left alone. Short interior gaps may be filled by interpolation, long ones are not.

For a gap of \(n\) samples between valid values \(x_a\) (just before) and \(x_b\) (just after), linear interpolation gives sample \(k\) of the gap, \(k = 1, \ldots, n\):

\[ x_k = x_a + \left(x_b - x_a\right)\frac{k}{n + 1} \]

For a resistivity use the same expression on \(\log_{10} x\).

Single-value calculator

No calculator: this method is a procedure, not a single equation.

Behavior

Before cleaning, the sentinel dominates the arithmetic: the mean of the 60-sample density curve is -331.413 g/cc and the mean of the neutron curve is -49.779. After conversion the density mean is 2.506 g/cc, with 20 of 60 samples missing, and the neutron mean is 0.180, with 3 of 60 samples missing. The neutron curve used -999.0 where the header said -999.25, which is why the header value alone is not enough. The gap report separates a leading gap of 8 samples, an interior gap of 2, an interior gap of 8 and a trailing gap of 2. With a limit of 3 samples, only the interior gap of 2 is filled (values 2.421 and 2.436 g/cc), and the 8-sample gap is left open.

Parameter guidance

Header null. Take it from the header, then verify it appears in the data. A header that says -999.25 over data that uses -999 is common.

Conventional list and minimum repeats. The list above is a starting point; add the values your own files use. Three repeats is a reasonable lower bound for a conventional value, though the header value itself needs only one.

Tolerance. Match to a small absolute tolerance (about 1e-6 here), because the sentinel may be rounded differently by different writers.

Out-of-range values. Use the plausible range of the family (see the step page). Be generous: the purpose is to catch non-physical numbers, not to judge the data. Zero is a special case: it is sometimes written in place of null in old or digitised logs, but it is also a real value for some curves. Treat it as null only when the whole curve segment is exactly zero.

Fill policy. Fill only short interior gaps, and set the maximum length from the sampling step and the vertical resolution of the tool: a gap shorter than the tool's own resolution carries almost no information that is lost. Longer gaps are better repaired by methods that use other curves, as covered under the washout and repair topics. Record which samples were filled, so that they are not mistaken for measurements.

Worked example

A synthetic density and neutron curve of 60 samples each, with a leading and a trailing null, two bad samples, an eight-sample gap, and a second null convention on the neutron curve. The code finds the sentinels, converts them, reports the gaps and fills the short ones:

import numpy as np

# Values that are commonly used to mean "no data". The LAS header NULL entry comes first, when it is given.
COMMON_SENTINELS = [-999.25, -999.0, -9999.0, -99999.0, 9999.0, 999.25, 1e30, -1e30]


def find_sentinels(x, header_null=None, min_count=3):
    """Exact repeated values that match a known sentinel, plus the header value."""
    x = np.asarray(x, dtype=float)
    found = set()
    for s in ([header_null] if header_null is not None else []) + COMMON_SENTINELS:
        if np.sum(np.isclose(x, s, rtol=0, atol=1e-6)) >= (1 if s == header_null else min_count):
            found.add(s)
    return sorted(found)


def to_nan(x, sentinels):
    x = np.array(x, dtype=float)
    for s in sentinels:
        x[np.isclose(x, s, rtol=0, atol=1e-6)] = np.nan
    return x


def gap_report(x):
    """Classify NaN runs as leading, trailing or interior."""
    bad = np.isnan(x)
    runs, i = [], 0
    while i < len(x):
        if bad[i]:
            j = i
            while j < len(x) and bad[j]:
                j += 1
            kind = "leading" if i == 0 else "trailing" if j == len(x) else "interior"
            runs.append((kind, i, j - i))
            i = j
        else:
            i += 1
    return runs


def fill_short_gaps(x, max_len=3, log=False):
    """Linear (or log-linear for resistivity) interpolation, only across interior gaps of max_len samples or fewer."""
    y = x.copy()
    for kind, start, n in gap_report(x):
        if kind == "interior" and n <= max_len:
            a, b = x[start - 1], x[start + n]
            t = np.arange(1, n + 1) / (n + 1)
            y[start:start + n] = np.exp(np.log(a) * (1 - t) + np.log(b) * t) if log else a * (1 - t) + b * t
    return y


if __name__ in ("__main__", "worked_example"):
    rng = np.random.default_rng(2)
    rhob = 2.5 + 0.08 * np.sin(np.linspace(0, 10, 60)) + rng.normal(0, 0.01, 60)
    rhob[:8] = -999.25            # above the logged interval
    rhob[30:32] = -999.25         # two bad samples inside the interval
    rhob[40:48] = -999.25         # an eight-sample gap
    rhob[58:] = -999.25           # below total depth
    nphi = 0.18 + rng.normal(0, 0.01, 60)
    nphi[10:13] = -999.0          # a second file convention for the same thing
    print("raw mean RHOB (nulls included):", round(float(rhob.mean()), 3), "g/cc")
    print("raw mean NPHI (nulls included):", round(float(nphi.mean()), 3))
    for name, x, header in (("RHOB", rhob, -999.25), ("NPHI", nphi, -999.25)):
        sent = find_sentinels(x, header)
        y = to_nan(x, sent)
        print(f"{name}: sentinels found {sent}; {int(np.isnan(y).sum())} of {len(y)} samples set to NaN")
        print(f"      mean after cleaning: {np.nanmean(y):.3f}; gaps: {gap_report(y)}")
    clean = to_nan(rhob, [-999.25])
    filled = fill_short_gaps(clean, max_len=3)
    print("after filling gaps of 3 samples or fewer:", gap_report(filled))
    print("filled values:", np.round(filled[30:32], 3), "(samples 30 and 31)")

Output

raw mean RHOB (nulls included): -331.413 g/cc
raw mean NPHI (nulls included): -49.779
RHOB: sentinels found [-999.25]; 20 of 60 samples set to NaN
      mean after cleaning: 2.506; gaps: [('leading', 0, 8), ('interior', 30, 2), ('interior', 40, 8), ('trailing', 58, 2)]
NPHI: sentinels found [-999.0]; 3 of 60 samples set to NaN
      mean after cleaning: 0.180; gaps: [('interior', 10, 3)]
after filling gaps of 3 samples or fewer: [('leading', 0, 8), ('interior', 40, 8), ('trailing', 58, 2)]
filled values: [2.421 2.436] (samples 30 and 31)

Assumptions and limitations

  • A sentinel value is never a real measurement. This holds for the standard values, but a curve can legitimately read 0 or -999.25 in a rare, non-physical way.
  • Leading and trailing gaps are structural (the curve was not run there) and not an error to repair.
  • A short interior gap is a lost sample in a smooth section, so interpolation is reasonable. Where a gap falls on a bed boundary, interpolation invents a gradient that is not there.
  • Every curve of a file uses the file's null value, or one of the conventional ones. A file may have several conventions after merging.

QC checks

  • No sentinel value is left in any curve. A search for exact repeats of the conventional values returns nothing.
  • The count of missing samples per curve is reasonable and matches the logged interval. A curve missing in the middle of the logged interval needs an explanation.
  • Statistics (mean, minimum, maximum) of each curve are in the plausible range of the family after conversion.
  • Filled samples are recorded and are a small fraction of the curve.
  • Calculations downstream propagate a missing value to a missing result, and do not turn it into zero.

Going Deeper

The LAS standard defines a null entry in the header, and -999.25 became the conventional value, but the standard cannot enforce that writers use it consistently, and data converted from other formats often carries other conventions. Missing values also matter for software: some systems treat NaN as missing in arithmetic and some do not, and a comparison against NaN is false, which can quietly exclude samples from a cutoff or a count. For that reason the conversion belongs at the data-loading boundary, once, and every later step then works with a single representation of a missing value. Gap filling is a modelling choice and not a data fix, so where it is used it should be visible in the result.

References

  1. Canadian Well Logging Society, 1992 (revised 1999). LAS Version 2.0: A Digital Standard for Logs. Update February 1992 (CWLS Log ASCII Standard).

Python reference implementation

Python reference implementation

The Python reference implementation is available to registered users with a verified email address. Register or sign in to view it.