Log Curve Cleanup
On this page
Purpose
A raw log is a measurement made in a hole, written by a tool, in a file, in one of several unit systems. Before any petrophysical calculation, it has to be made into a clean, consistent curve: in known units, free of placeholder values, on a trustworthy depth scale, without non-geological spikes, and, for the neutron, corrected with the right tool-specific method. Log curve cleanup is that step. It does not repair bad data from a washout (see the washout topic) and does not shift the level of a curve (see normalization). It makes sure that what comes out of the file is what the tool measured, in a form the equations expect.
The errors it prevents are large and silent: a density in kg/m³ read as g/cm³ gives a porosity of about -1500, a -999.25 left in a mean pulls it hundreds of units away, and a 1.5 ft depth error puts a perfectly good log against the wrong bed.
Position in the workflow
Upstream. Curve Aliasing and Mnemonics says which curve is which. Cleanup depends on it: the unit tests, the range checks and the sentinel conventions are per family.
Downstream. Everything. The cleaned curves feed:
- normalization and the washout topic, which assume sentinels are gone, units are right and depth is consistent,
- every Stage 2 calculation, from clay volume to saturation, and
- the neutron porosity used with the density in the porosity step, which depends on the correction chosen here.
Error propagation. A unit error is multiplicative and obvious in a plot, but only if someone looks. A null value is enormous and visible in a histogram. A depth error is small and subtle and goes straight into the thin-bed answer. A wrong neutron correction is a few porosity units, which is the size of the signal in a tight reservoir.
Key concepts
Units and conversion. Curves come in imperial and SI forms (g/cm³ and kg/m³, µs/ft and µs/m, in and mm, v/v and percent, ohm·m and mS/m). The conversion is simple, and the detection is the work. See Unit Detection and Conversion.
Null values. A Null sentinel value is a flag and not a measurement. It must be turned into a true missing value, once, at loading. See Null and Invalid Value Handling.
Depth. The index must be monotonic, regular and on a stated Depth reference and unit. Runs and core have depth shifts between them. See Depth Reference and Sampling Checks.
Spikes. Narrow excursions that are not geology. A Median filter or Hampel filter removes them, at the cost of thin beds. See Spike Removal and Filtering.
Neutron tool. The neutron log depends on the tool, the hole and the matrix, so it is the one curve that cannot be used as delivered without knowing how it was made. See Neutron Tool Selection.
Order matters. Statistics need nulls removed, unit detection needs statistics, filters need uniform depth, and neutron checks need clean density and caliper.
Method selection guide
| Method | Inputs | Use when | Strengths | Weaknesses |
|---|---|---|---|---|
| Unit detection and conversion | Curve values, header units, family | On every curve of every file | Exact factors, simple tests that work for density, caliper and porosity | Sonic and conductivity are ambiguous by value alone and need the header or a cross-check |
| Null and invalid value handling | Curve values, header null, plausible ranges | First, on every curve | Removes the largest errors, cheap, transparent | Fill rules are a judgement; a sentinel that is also a real value is missed |
| Depth reference and sampling checks | Depth index, elevations, reference curve | Before splicing, normalization or any comparison with core or tops | Catches errors that affect every curve | A constant shift cannot fix stretch; large shifts often mean a reference error |
| Spike removal and filtering | One curve, window, threshold | Isolated spikes and noise in an otherwise good curve | Hampel changes only outliers and keeps edges | Removes real thin beds narrower than half the window; cannot repair wide artefacts |
| Neutron tool selection | Tool, size, scale, hole fluid, salinity, caliper, density | Before neutron porosity is used in a calculation | Removes a tool- and hole-dependent bias | Needs facts that are often missing; no universal chart |
Decision guidance
- Always do the first three, in order, on every well. They are not optional.
- Use a filter only for a cause you can name (a telemetry spike, noise). Do not smooth a curve to make it look better.
- Choose a Hampel filter over a plain median when most samples are good, and a plain median only when the curve is mostly noise.
- Treat the neutron separately: do not use it in a porosity or clay calculation until its tool and correction are known, or its uncertainty is stated.
- If a check fails and you cannot say why, stop and ask a person. Automatic repair of an unexplained error hides it.
Shared parameter picking
Null values. The header value, the conventional list, the minimum repeat count and the tolerance. One list for the project.
Plausible range per family. Used to detect units, to find invalid samples and for QC. Keep it in one place, shared with Curve Aliasing and Mnemonics, and make it wide.
Unit detection thresholds. The percentile limits for sonic, the median limits for density and caliper, and the percent limit for porosity. Start from the values on the unit page and adjust to your data.
Depth tolerances. The tolerance on the step, the search window for shifts, and the elevations of the depth references.
Filter settings. Window in feet (not in samples), threshold k, the noise floor for each curve, and the maximum fraction of a curve that may be altered before it is sent to review.
Neutron facts. Tool family and model, diameter, lithology scale, mud and salinity for each well or run. Record them once and store them with the well.
Recommended default approach
Absent other information, a careful generalist would:
- Read the header: units, null value, depth reference and elevations, tool and run information.
- Convert null sentinels to missing values, and report the gaps. Fill only short interior gaps, and record them.
- Check the depth index (units, step, order, duplicates, gaps) and put every data set on one depth reference.
- Detect the unit of each curve from header and values, convert to the working units (g/cm³, µs/ft, in, v/v, ohm·m), and list every curve where the two disagree.
- Despike with a Hampel filter, a window of 5 to 7 samples and k = 3, only the curves that need it. Check the fraction flagged and that thin beds survive.
- Collect the neutron tool facts, apply the matching correction or state that the curve is used as delivered, and check it against density in a clean interval.
- Write the record of everything changed, and keep the raw curves.
Combining methods
These methods are steps in sequence and not alternatives. Run null handling before unit detection, because the statistics must be free of sentinels. Run the depth checks before any filter that assumes uniform sampling. Despike after unit conversion so that thresholds are in working units. Check the neutron last, because its check needs the density, the caliper and the depth to be clean. Filtering and gap filling are the only steps that change values rather than labels, so keep their results separate from the raw data.
QC of results
A good result:
- has every curve in the working units, with the unit label updated,
- has no sentinel values, with all missing samples explained,
- has a single, monotonic, regular depth index on one stated reference,
- has few flagged spikes, mostly isolated, with thin beds intact, and
- has a neutron curve whose tool, scale and correction are stated and that agrees with density in clean limestone.
Signs of a bad result: porosity values in the hundreds or thousands, negative densities, a mean that is far from the median, a step at every run boundary, a neutron that follows the caliper, and a filtered curve that has lost a marker bed.
Common pitfalls
- Computing a statistic before removing the nulls, and then using it to detect units.
- Trusting the unit string in the header, or ignoring it.
- Detecting the sonic unit from values in a narrow interval, where the two ranges overlap.
- Treating a conductivity as a resistivity, or averaging resistivity linearly.
- Applying a unit conversion twice because the label was not changed.
- Mixing depth references between logs, core and tops.
- Despiking with a window wider than the thin beds that matter, or despiking the caliper before using it as a flag.
- Filling long gaps by interpolation and then treating the result as data.
- Using a neutron correction chart for the wrong tool diameter, or applying a correction to a curve that is already corrected.
- Reporting the neutron without its lithology scale.
Going Deeper
Cleanup is unglamorous and the most error-prone part of any interpretation, because it combines many small facts about files, tools and conventions, and because its errors look like geology. Automating it is attractive, and the checks on these pages are all automatable, but the decisions at the margin (an ambiguous sonic unit, a long gap, a spike that may be a bed) need a person who can see the log and the well history. Good practice is to automate the detection and the record-keeping, and to leave the ambiguous cases for review. A second theme is provenance: each step should say what it changed and why, so that a later reader can undo it. Standards for log data formats improve the situation by carrying units and descriptions with the data, but the quality of what is written in them is still the operator's and the vendor's.