CamPetro

Deterministic Rules from Crossplots

On this page

Summary

Deterministic lithofacies rules assign each depth sample to a facies with an ordered table of thresholds read from crossplots: gamma ray separates shale from clean rock, the Neutron-density separation separates sandstone from carbonate, and the bulk density separates limestone from dolomite. The rules are transparent and repeatable, need no training data, and make a baseline against which other methods can be judged. Use them when the logs are good, the lithologies are few, and the thresholds can be tied to known rock.

Inputs and outputs

Item Units
Input Gamma ray, Bulk density, Neutron porosity (cleaned and normalized); optionally True formation resistivity and Photoelectric factor gAPI, g/cm³, v/v, ohm·m
Derived Neutron-density separation v/v
Output Facies code (one integer per depth sample, null where an input is null) none

Equations

The neutron-density separation compares the neutron porosity with the density porosity, both on the same matrix scale. On the limestone scale, with matrix density 2.71 g/cm³ and fluid density 1.0 g/cm³,

\[ \Delta\phi_{ND} = \phi_N - \frac{2.71 - \rho_b}{2.71 - 1.0} \]

It is positive in shale, negative in clean sandstone and close to zero in limestone. The rule table is evaluated top to bottom and the first rule that matches wins:

Order Rule Facies
1 any of GR, RHOB or NPHI is null null (no result)
2 GR ≥ 75 gAPI Shale
3 GR < 75 and ΔφND < -0.01 Sandstone
4 GR < 75, ΔφND ≥ -0.01 and RHOB ≥ 2.66 g/cm³ Dolomite
5 GR < 75 and ΔφND ≥ -0.01 Limestone

The thresholds are placed between the typical values of neighbouring facies, and the table is only as good as those placements. The accuracy is judged with a confusion matrix, whose entry \(C_{ij}\) counts the samples of true facies \(i\) that were given facies \(j\):

\[ \text{accuracy} = \frac{\sum_i C_{ii}}{\sum_{ij} C_{ij}} \qquad \text{recall}_i = \frac{C_{ii}}{\sum_j C_{ij}} \qquad \text{precision}_j = \frac{C_{jj}}{\sum_i C_{ij}} \]

Single-value calculator

No calculator: this method is a procedure, not a single equation. The worked example below runs the whole procedure on synthetic logs.

Behavior

The synthetic well has 687 samples in four facies. The neutron-density separation averages +0.193 in shale, -0.056 in sandstone, +0.012 in limestone and +0.095 in dolomite, each with a spread near 0.045. Shale is the easy case, recovered with 97% recall because gamma ray and separation both point the same way. The hard case is the limestone against the sandstone: the separations differ by only 0.07 while the spread is 0.05, so 48 of 318 sandstone samples are called limestone and 47 of 122 limestone samples are called sandstone. The overall accuracy is 0.837, but the recall of limestone is only 0.59. A single overall number hides this: read the matrix, not the accuracy.

Parameter guidance

Thresholds. Place each one between the facies it separates, using crossplots of the key well, core descriptions and the known matrix values: clean sandstone near 2.65 g/cm³, limestone near 2.71 and dolomite near 2.85 to 2.87, with the porosity lowering all of them. Moving a threshold trades recall of one facies for another, so tune it against core-described intervals and not by eye alone. Matrix and fluid for the separation. Use the matrix and fluid values of the neutron scale of the tool, normally limestone; see Neutron Tool Selection for scale and tool differences. Resistivity and photoelectric factor. Extra rules can use them, for example a high resistivity with high gamma ray for an organic-rich shale, or the photoelectric factor to separate dolomite from limestone. Resistivity responds to fluid as well as rock, so apply it last and only to a shale branch. Inputs. Rules assume normalized, repaired logs: see Normalization.

Worked example

A synthetic well of 40 beds in four facies, a rule table written as a function, and the confusion matrix computed by hand against the known facies. The code is plain numpy and runs with a fixed seed:

import numpy as np

NAMES = ["Shale", "Sandstone", "Limestone", "Dolomite"]
# facies means: GR (API), RHOB (g/cm3), NPHI (v/v, limestone scale), Rt (ohm.m). Illustrative values.
MEAN = np.array([[105.0, 2.55, 0.30, 3.0],
                 [45.0, 2.34, 0.17, 12.0],
                 [28.0, 2.60, 0.09, 35.0],
                 [28.0, 2.72, 0.07, 45.0]])


def synthetic_well(n_beds=40, seed=7, probs=(0.35, 0.30, 0.25, 0.10), bed_scale=1.0):
    """Beds of 8 to 25 samples. Each bed has its own offset, shared by all its samples, and each
    sample adds independent noise. bed_scale multiplies the bed offsets. Resistivity is simulated in log10."""
    rng = np.random.default_rng(seed)
    truth, logs = [], []
    for _ in range(n_beds):
        f = rng.choice(4, p=probs)
        n = rng.integers(8, 26)
        bed = rng.normal(0, np.array([8.0, 0.04, 0.025, 0.20]) * bed_scale)
        gr = MEAN[f, 0] + bed[0] + rng.normal(0, 12.0, n)
        rhob = MEAN[f, 1] + bed[1] + rng.normal(0, 0.04, n)
        nphi = MEAN[f, 2] + bed[2] + rng.normal(0, 0.025, n)
        rt = 10 ** (np.log10(MEAN[f, 3]) + bed[3] + rng.normal(0, 0.15, n))
        truth += [f] * n
        logs.append(np.column_stack([gr, rhob, nphi, rt]))
    return np.vstack(logs), np.array(truth)


def nd_separation(rhob, nphi, rho_ma=2.71, rho_fl=1.0):
    """Neutron minus density porosity, both on the limestone scale."""
    return nphi - (rho_ma - rhob) / (rho_ma - rho_fl)


def classify(gr, rhob, nphi):
    """Ordered rule table: the first rule that matches wins. A missing input gives -1 (no result)."""
    ndsep = nd_separation(rhob, nphi)
    out = np.zeros(gr.shape, dtype=int)                 # default: shale
    clean = gr < 75.0
    out[clean & (ndsep < -0.01)] = 1                    # clean, neutron below density: sandstone
    out[clean & (ndsep >= -0.01)] = 2                   # clean, no negative separation: limestone
    out[clean & (ndsep >= -0.01) & (rhob >= 2.66)] = 3  # clean and dense: dolomite
    out[np.isnan(rhob) | np.isnan(nphi) | np.isnan(gr)] = -1
    return out


logs, truth = synthetic_well()
gr, rhob, nphi, rt = logs.T
pred = classify(gr, rhob, nphi)

print(f"{len(truth)} samples; limestone-scale ND separation by true facies:")
sep = nd_separation(rhob, nphi)
for k, name in enumerate(NAMES):
    s = sep[truth == k]
    print(f"  {name:10s} n={len(s):4d}  mean {s.mean():+.3f}  sd {s.std():.3f}")

# confusion matrix by hand: rows are true facies, columns are predicted facies
cm = np.zeros((4, 4), dtype=int)
np.add.at(cm, (truth, pred), 1)
print("\nconfusion matrix (rows true, columns predicted):")
print("            " + " ".join(f"{n[:5]:>6s}" for n in NAMES))
for k, name in enumerate(NAMES):
    print(f"  {name:10s}" + " ".join(f"{v:6d}" for v in cm[k]))
print(f"\noverall accuracy {np.trace(cm) / cm.sum():.3f}")
for k, name in enumerate(NAMES):
    rec = cm[k, k] / cm[k].sum()
    prec = cm[k, k] / max(cm[:, k].sum(), 1)
    print(f"  {name:10s} recall {rec:.3f}  precision {prec:.3f}")

Output

687 samples; limestone-scale ND separation by true facies:
  Shale      n= 191  mean +0.193  sd 0.044
  Sandstone  n= 318  mean -0.056  sd 0.044
  Limestone  n= 122  mean +0.012  sd 0.052
  Dolomite   n=  56  mean +0.095  sd 0.043

confusion matrix (rows true, columns predicted):
             Shale  Sands  Limes  Dolom
  Shale        186      0      5      0
  Sandstone      4    266     48      0
  Limestone      0     47     72      3
  Dolomite       0      0      5     51

overall accuracy 0.837
  Shale      recall 0.974  precision 0.979
  Sandstone  recall 0.836  precision 0.850
  Limestone  recall 0.590  precision 0.554
  Dolomite   recall 0.911  precision 0.944

Assumptions and limitations

  • Each facies has a distinct, stable position on the logs, so that thresholds fixed in advance separate them. Porosity, fluid and mineral mixtures move a sample across a threshold.
  • The logs are normalized, in the same units and tool scale, in every well in which the table is applied. A shifted density log moves every sample across a threshold.
  • The neutron-density separation is meaningful, which needs a good density log and a neutron log on the scale the matrix and fluid values assume. Washouts, gas and heavy minerals break it.
  • The facies set is small and defined by the logs. A facies defined by texture or by depositional environment cannot be separated by thresholds on these logs.
  • The first matching rule is the right one. Overlapping rules are resolved only by their order, which is a judgement.

QC checks

  • Every depth sample receives exactly one code or null, and null appears only where an input is null.
  • The rule thresholds sit in gaps between the clusters on the key crossplots. A threshold in the middle of a cloud of points is a sign of a bad rule.
  • The result against core description is shown as a confusion matrix, with recall and precision by facies, and not as one accuracy number.
  • The facies track is thicker than a few samples. A track that flickers every few samples shows that the samples sit near a threshold: consider a depth filter or a wider gap.
  • The proportions of each facies in the well are consistent with the geology. A sandstone fraction that doubles in a carbonate section means the separation or the density has a problem.

Going Deeper

Rule tables are the oldest form of automated lithology: they are the crossplot interpretation written down. Their strength is that any geologist can read, challenge and adjust them, and their weakness is that thresholds fixed by hand fit the well they were drawn on. A rule table is a good baseline and a good first label set for the other methods: clustering can be labelled with the same rules, and the table can provide training labels where no core description exists, with the caution that a classifier trained on its own rules only reproduces them. Rules can be made softer by replacing a hard threshold with a membership function, which gives a degree of belief for each facies and identifies the samples that sit close to a boundary. Pre-defined thresholds for evaporites and casing, and for coal, are often put in front of the table because their densities are extreme and unambiguous.

References

References will be added once verified.

Python reference implementation

Python reference implementation

The Python reference implementation is available to registered users with a verified email address. Register or sign in to view it.