CamPetro

Histogram and Percentile-Based Picks

On this page

Summary

A normalization is only as good as the percentile picks that anchor it. This page covers how a pick is computed, which percentiles to use, how thick a Normalization reference interval needs to be, and why a percentile is safer than the minimum or maximum. It is a procedure, not a single calculation.

Inputs and outputs

Item Units
Input Curve samples in a reference interval, one well curve units
Input Percentile levels for the low, middle and high picks %
Output Low, middle and high picks (Source low pick, Source middle pick, Source high pick) curve units
Output A decision on whether the interval is thick and clean enough

Equations

For \(n\) samples sorted into ascending order \(x_{(1)} \le \dots \le x_{(n)}\), the \(\nrmPct\)-th percentile by linear interpolation is

\[ h = (n-1)\,\frac{\nrmPct}{100}, \qquad p = x_{(\lfloor h\rfloor + 1)} + \left(h - \lfloor h\rfloor\right)\left(x_{(\lfloor h\rfloor + 2)} - x_{(\lfloor h\rfloor + 1)}\right) \]

with \(x_{(n+1)}\) taken equal to \(x_{(n)}\). This is the default of most numerical libraries; other conventions differ by a fraction of a sample spacing and are unimportant here. The low, middle and high picks are this value at three levels: \(\nrmSrcLow\), \(\nrmSrcMid\) and \(\nrmSrcHigh\).

For independent samples from a distribution with density \(f\) the scatter of a pick shrinks as the number of samples grows:

\[ \mathrm{sd}(p) \approx \frac{\sqrt{\nrmPct\,(100-\nrmPct)}\,/\,100}{f(p)\,\sqrt{n}} \]

Log samples are strongly autocorrelated, so the effective number of independent samples is much smaller than the count of depth steps. The formula therefore gives a lower bound on the scatter, and the thickness test in the worked example is the practical guide.

Symbol Variable Units Typical range
\(q\) Percentile level % 2 to 98
\(p_{\mathrm{lo}}\) Source low pick
\(p_{\mathrm{mid}}\) Source middle pick
\(p_{\mathrm{hi}}\) Source high pick
Normalization reference interval
Key well

Single-value calculator

No calculator: picking percentiles is a procedure that depends on the shape of the distribution in your interval, not a single equation. The worked example computes the quantities on a synthetic interval.

Behavior

A percentile is robust to outliers where a minimum or maximum is not. In the worked example a single 290 gAPI spike in 200 samples moves the maximum from 130.4 to 290.0 but moves the 95th percentile only from 122.9 to 123.1. The pick is also less repeatable in a thin interval: the scatter of the 95th percentile across repeated draws falls from 3.5 gAPI for a 5 ft interval to 1.6 gAPI at 50 ft and 1.1 gAPI at 100 ft. These are the numbers for a uniform shale with a standard deviation of 8 gAPI and independent samples. Real logs are autocorrelated and layered, so the real scatter will be larger and the benefit of thickness smaller. With three spikes in 50 ft the 95th percentile scatters by 2.1 gAPI against 21.4 for the maximum.

Parameter guidance

Which percentiles. The usual picks are the 5th and 95th for the low and high ends and the 50th for the middle. They are far enough out to capture the spread of the curve and far enough in to ignore the worst bad-hole and spike samples. For a thick, clean interval with good data, the 2nd and 98th use more of the distribution. For a short or noisy interval, move in to the 10th and 90th. The 1st/99th and the minimum/maximum are not robust and are not recommended.

What the picks should represent. For a gamma ray the low pick should represent clean rock and the high pick shale, so choose the interval to contain both. For a curve where only the shale is stable (density or neutron in a regional shale), take a narrow percentile range (such as 25th and 75th) inside that one lithology, because the extremes then contain a mix of tool noise and a few genuine beds.

Interval thickness. Use the thickest laterally consistent interval available. As a working rule, tens of feet of uniform rock (at least 100 samples at a typical 0.5 ft step) is the minimum for a stable 5th or 95th percentile; the worked example shows how the scatter falls with thickness. A thicker interval that crosses a geological change is worse than a thinner uniform one.

Reading the histogram. Plot the histogram of the curve in the interval for both wells on the same axes before picking. A single sharp mode means the interval is a uniform lithology; two modes (clean and shale) mean the low and high percentiles will fall in different populations, which is what a two-point map wants. Verify the picks by drawing them on the histogram and checking they sit where you expect. A pick that falls in the gap between two modes is unstable: a small change in the interval moves it a long way.

Cleaning before picking. Remove nulls, flagged bad hole and spikes first (see Washout Identification and Repair), then take the picks. Picking before cleaning bakes the damage into the normalization. The way the picks are used is on Shift, Scale, and Shift-and-Scale.

Worked example

Four checks on synthetic shale intervals with a mean of 110 gAPI and a standard deviation of 8 gAPI: the definition of a percentile worked by hand, the effect of a single spike, the repeatability of the 95th percentile against interval thickness at 0.5 ft sampling, and the percentile against the maximum when a few spikes are present.

rng = np.random.default_rng(3)

def interval(n, spikes=0):
    """Synthetic shale reference interval, GR ~ N(110, 8), with optional spikes."""
    g = rng.normal(110.0, 8.0, n)
    if spikes:
        idx = rng.choice(n, spikes, replace=False)
        g[idx] = rng.uniform(190, 300, spikes)      # bad-sample spikes
    return g

# 1. Percentile definition: linear interpolation between order statistics
x = np.sort(interval(11))
h = (len(x) - 1) * 0.90
lo = int(np.floor(h))
p90 = x[lo] + (h - lo) * (x[lo + 1] - x[lo])
print(f"P90 by hand = {p90:.2f}, np.percentile = {np.percentile(x, 90):.2f}")

# 2. Robustness: one spike in 200 samples
clean = interval(200)
spiked = clean.copy()
spiked[17] = 290.0
print("\none spike of 290 in 200 samples")
print(f"            clean    spiked")
print(f"max       {clean.max():7.1f}  {spiked.max():7.1f}")
print(f"P95       {np.percentile(clean, 95):7.1f}  {np.percentile(spiked, 95):7.1f}")
print(f"P99       {np.percentile(clean, 99):7.1f}  {np.percentile(spiked, 99):7.1f}")

# 3. Repeatability of a pick versus interval thickness (0.5 ft sampling)
print("\nscatter of the P95 pick over 400 repeated intervals")
print(f"{'thickness':>10} {'samples':>8} {'sd of P95':>10}")
for ft in (5, 10, 25, 50, 100):
    n = int(ft / 0.5)
    picks = [np.percentile(interval(n), 95) for _ in range(400)]
    print(f"{ft:8d} ft {n:8d} {np.std(picks):10.2f}")

# 4. The same intervals with a few spikes: sd of the pick for P95 versus the maximum
picks95 = [np.percentile(interval(100, 3), 95) for _ in range(400)]
picksmx = [interval(100, 3).max() for _ in range(400)]
print(f"\n50 ft with 3 spikes: sd of P95 = {np.std(picks95):.1f}, sd of maximum = {np.std(picksmx):.1f}")

Output

P90 by hand = 126.33, np.percentile = 126.33

one spike of 290 in 200 samples
            clean    spiked
max         130.4    290.0
P95         122.9    123.1
P99         127.4    129.7

scatter of the P95 pick over 400 repeated intervals
 thickness  samples  sd of P95
       5 ft       10       3.54
      10 ft       20       3.20
      25 ft       50       2.33
      50 ft      100       1.59
     100 ft      200       1.11

50 ft with 3 spikes: sd of P95 = 2.1, sd of maximum = 21.4

Assumptions and limitations

  • The interval is lithologically uniform (or uniformly mixed) in both wells, so that the same percentiles describe the same rock.
  • Samples are free of nulls, flagged bad hole and spikes, or enough of them have been removed that the percentiles are not driven by them.
  • Samples are treated as independent in the scatter formula. They are not; the formula is optimistic.
  • The distribution is stable over the interval. A trend with depth (compaction, a change of mud) shifts the picks with the interval chosen.

QC checks

  • Overlay the histograms of both wells before and after normalization.
  • Re-pick with the interval shifted by 10 to 20 percent of its thickness. The picks should move by much less than the offset being corrected.
  • Compare picks at two or three percentile levels (for example the 5th/95th and 10th/90th). If the normalization changes materially, the picks are unstable.
  • Check that no pick sits on an isolated spike or in an empty part of the histogram.
  • Record the interval top and base, the percentile levels and the sample count with each normalization.

Going Deeper

The percentile is the simplest robust estimate of a distribution's end-points. More elaborate methods, such as picking the mode of a histogram with kernel smoothing or matching the full quantile function, can reduce the sensitivity to the choice of percentile, but they need more data and are harder to audit. The recurring difficulty is that the 'same rock' in two wells is an assumption. Percentile picks make the assumption explicit and testable (by comparing the distributions) but cannot validate it. A correction that leaves the distributions matching in a reference interval is not proof that the curves are correct outside it.

References

  1. Shier, D.E., 2004. Well log normalization: methods and guidelines. Petrophysics, 45(3), 268–280.
  2. Neinast, G.S. and Knox, C.C., 1973. Normalization of well log data. Transactions of the SPWLA 14th Annual Logging Symposium, Paper I.

Python reference implementation

Python reference implementation

The Python reference implementation is available to registered users with a verified email address. Register or sign in to view it.