CamPetro

Repair: Machine-Learning Infill

On this page

Summary

Machine-learning infill fills a flagged interval with values predicted by a model, such as a random forest, trained on other curves in the same well. It is prediction, not measurement: it can look plausible while being wrong, it cannot see what the other curves do not, and its apparent accuracy depends on how it is validated. It is a convenience for continuous curves, not a way to recover the log that was lost.

Inputs and outputs

Item Units
Input Undamaged curves from the same well (predictors), and the damaged target curve over good samples (Bulk density) curve units
Input A Bad-hole flag and padding that define the interval to infill 0 or 1
Output A Machine-learning infill curve over the flagged interval, and a flag showing where it was used curve units; 0 or 1

Equations

A random forest is the mean of \(B\) regression trees, each grown on a bootstrap sample of the training data with a random subset of the predictors at each split:

\[ \hat y(\mathbf x) = \frac{1}{B}\sum_{b=1}^{B} T_b(\mathbf x) \]

Each tree is a piecewise constant function, so the forest's prediction is always an average of training values. It therefore stays inside the range of the training targets and cannot extrapolate.

The skill of a model is measured on samples it has not seen. For held-out samples \(i\) with true value \(\rhob_i\) and prediction \(\hat\rho_i\):

\[ \mathrm{RMSE} = \sqrt{\frac{1}{n}\sum_{i}\left(\hat\rho_i - \rhob_i\right)^2}, \qquad R^2 = 1 - \frac{\sum_i \left(\hat\rho_i - \rhob_i\right)^2}{\sum_i \left(\rhob_i - \overline{\rhob}\right)^2} \]

The held-out samples must be separated from the training samples in depth. Samples at neighbouring depths are highly correlated, so a random split places near-copies of every test sample in the training set, and the score measures interpolation between neighbours and not the ability to predict an unseen interval. This is Data leakage. A depth-block split holds out whole intervals at least as long as the gap that will be infilled.

Symbol Variable Units Typical range
\(\rho_b\) Bulk density g/cm³ 1.8 to 3.0
Machine-learning infill
Data leakage
Bad hole

Single-value calculator

No calculator: a machine-learning model is a procedure with a training set, not a single equation. The worked example trains a small random forest in numpy on a synthetic well.

Behavior

The honest result of a machine-learning infill is often unremarkable. In the worked example a random forest and an ordinary linear regression give almost the same score on a synthetic well. With a random split of the samples the forest has an R² of 0.844 (RMSE 0.055 g/cm³) and the linear model 0.831 (0.057); with depth blocks of 100 ft the two are the same, 0.794 (0.063 and 0.064 g/cm³). The random split overstates the skill of the forest (0.844 against 0.794) because neighbouring samples leak. In a 30 ft gap the forest recovers the density with an RMSE of 0.091 g/cm³ and a bias of -0.073 g/cm³: it misses the thin dense beds that the other curves barely show. The true spread in the gap is 0.102 g/cm³ (sd) and the predicted spread 0.076, and the true range is 2.19 to 2.72 against a predicted 2.15 to 2.47. The prediction is smoother than the truth and loses the extremes. For a coal-like input (GR 30, NPHI 0.55, DT 125) outside the training data, the forest returns 2.15 g/cm³, inside the training range of 2.04 to 2.79, where a linear model extrapolates to 1.60. Neither is a measurement, and the forest's answer looks reasonable and is not.

Parameter guidance

Whether to use it. Use it only where a continuous curve is needed for visual work, cross-plots, or a model that cannot take nulls, and mark the result. For a quantitative result such as porosity or saturation, prefer a null or a regression with a stated band, because a model that is trained on the other curves cannot add information to a calculation that uses those same curves.

Predictors. Use curves that the hole did not damage: gamma ray, deep resistivity and often the sonic. Do not use the neutron or PE from the same damaged interval, and do not use a curve that is itself a repair. Depth itself and position-based features make a model memorize the training interval and should be avoided.

Training data. Train only on good samples from the same well and formation, with the flagged and padded samples removed, near the gap. Exclude intervals with a known different rock (coal, salt) unless the gap is in one.

Validation. Hold out whole depth blocks, at least as long as the largest gap, and report the error on them, as in the example. A random split gives an optimistic number. A rough sense of repeatability comes from training on different blocks and seeing how the infill changes. The reference for the cross-validation of data with spatial or temporal structure is in the references below.

Model settings. The number of trees (100 to several hundred), the depth and the minimum leaf size matter less than the predictors and the validation. Use a standard library implementation in practice. The example is a small teaching version.

Infill only or replace. Infill only fills the samples that are already null, such as those removed by Repair: Null-Out: the original data are never touched and the infill is confined to the gap. Replace overwrites flagged samples that still have a value. It is more aggressive, and is only acceptable when the flagged samples are known to be bad. In both cases write the result to a new curve and keep the original, plus a flag curve for the predicted samples.

Uncertainty. Report the validation RMSE with the repaired curve. The spread among the trees of a forest is a measure of disagreement among the trees and not of the true error, which is usually larger.

Worked example

A synthetic 500 ft well in which each 100 ft formation has its own matrix-density offset and thin cemented beds are only weakly visible on the other curves. A small random forest and a linear regression are scored with a random split and with depth blocks, then used to infill a 30 ft gap whose true density is known, and finally given an input outside the training range.

rng = np.random.default_rng(12)
n = 1000                                         # 500 ft at 0.5 ft
smooth = lambda a, k: np.convolve(a, np.ones(k) / k, mode="same")

# --- Synthetic well. Each 100 ft "formation" has its own matrix density offset, as real wells do.
vsh = np.clip(0.5 + 3.0 * smooth(rng.normal(0, 1, n), 25), 0, 1)
phi = np.clip(0.28 - 0.18 * vsh + 0.06 * smooth(rng.normal(0, 1, n), 9), 0.02, 0.35)
cem = (smooth(rng.normal(0, 1, n), 5) > 0.55) * 1.0       # thin cemented beds, only weakly seen by the other logs
formation = np.arange(n) // 200
offset = rng.normal(0, 0.04, 5)[formation]
gr = 20 + 100 * vsh + rng.normal(0, 6, n)
nphi = phi + 0.12 * vsh - 0.03 * cem + rng.normal(0, 0.012, n)
dt = 56 + 120 * phi + 20 * vsh - 8 * cem + rng.normal(0, 2.5, n)
rhob = 2.65 - 1.65 * phi + 0.02 * vsh + 0.20 * cem + offset + rng.normal(0, 0.012, n)
X = np.column_stack([gr, nphi, dt])

# --- A very small random forest (bootstrap samples, random feature subsets, mean of trees)
def grow(X, y, rs, depth=0, max_depth=7, min_leaf=8):
    if depth == max_depth or len(y) < 2 * min_leaf:
        return float(y.mean())
    best = None
    for j in rs.choice(X.shape[1], 2, replace=False):
        for t in np.quantile(X[:, j], np.linspace(0.1, 0.9, 7)):
            left = X[:, j] <= t
            if left.sum() < min_leaf or (~left).sum() < min_leaf:
                continue
            sse = ((y[left] - y[left].mean()) ** 2).sum() + ((y[~left] - y[~left].mean()) ** 2).sum()
            if best is None or sse < best[0]:
                best = (sse, j, t, left)
    if best is None:
        return float(y.mean())
    _, j, t, left = best
    return (j, t, grow(X[left], y[left], rs, depth + 1), grow(X[~left], y[~left], rs, depth + 1))

def walk(tree, x):
    while isinstance(tree, tuple):
        tree = tree[2] if x[tree[0]] <= tree[1] else tree[3]
    return tree

def forest_fit(X, y, trees=12, seed=0):
    rs = np.random.default_rng(seed)
    out = []
    for _ in range(trees):
        idx = rs.integers(0, len(y), len(y))
        out.append(grow(X[idx], y[idx], rs))
    return out

def forest_predict(forest, X):
    return np.array([np.mean([walk(t, x) for t in forest]) for x in X])

def linear_fit_predict(Xtr, ytr, Xte):
    A = np.column_stack([np.ones(len(ytr)), Xtr])
    beta, *_ = np.linalg.lstsq(A, ytr, rcond=None)
    return np.column_stack([np.ones(len(Xte)), Xte]) @ beta

def r2(y, p):
    return 1 - ((y - p) ** 2).sum() / ((y - y.mean()) ** 2).sum()

# --- Validation: random split versus depth-block split (5 blocks of 100 ft)
perm = np.random.default_rng(1).permutation(n)
random_fold = np.empty(n, int)
random_fold[perm] = np.arange(n) % 5
block_fold = formation
for label, fold in (("random split", random_fold), ("depth blocks", block_fold)):
    rf_pred = np.empty(n)
    lin_pred = np.empty(n)
    for k in range(5):
        te = fold == k
        rf_pred[te] = forest_predict(forest_fit(X[~te], rhob[~te]), X[te])
        lin_pred[te] = linear_fit_predict(X[~te], rhob[~te], X[te])
    print(f"{label:13s} R2: random forest {r2(rhob, rf_pred):.3f}   linear {r2(rhob, lin_pred):.3f}   "
          f"RMSE (g/cm3): forest {np.sqrt(np.mean((rhob - rf_pred) ** 2)):.3f}  linear {np.sqrt(np.mean((rhob - lin_pred) ** 2)):.3f}")

# --- Infill of a washout gap: samples 420 to 480 are lost (60 samples = 30 ft), with 15 samples of padding
gap = np.arange(420, 480)
train = np.ones(n, bool)
train[405:495] = False
forest = forest_fit(X[train], rhob[train])
pred = forest_predict(forest, X[gap])
err = pred - rhob[gap]
print(f"\ninfill of the gap, compared with the (hidden) true values:")
print(f"  RMSE {np.sqrt(np.mean(err ** 2)):.3f} g/cm3, bias {err.mean():+.3f}")
print(f"  true spread in the gap sd = {rhob[gap].std():.3f}, predicted spread sd = {pred.std():.3f}")
print(f"  true range {rhob[gap].min():.2f} to {rhob[gap].max():.2f}; predicted range {pred.min():.2f} to {pred.max():.2f}")

# --- Extrapolation: features far outside the training data (a coal-like response)
coal_like = np.array([[30.0, 0.55, 125.0]])
print(f"\ncoal-like input (GR 30, NPHI 0.55, DT 125): training density range {rhob[train].min():.2f} to {rhob[train].max():.2f}, "
      f"forest returns {forest_predict(forest, coal_like)[0]:.2f}, linear returns {linear_fit_predict(X[train], rhob[train], coal_like)[0]:.2f}")

Output

random split  R2: random forest 0.844   linear 0.831   RMSE (g/cm3): forest 0.055  linear 0.057
depth blocks  R2: random forest 0.794   linear 0.794   RMSE (g/cm3): forest 0.063  linear 0.064

infill of the gap, compared with the (hidden) true values:
  RMSE 0.091 g/cm3, bias -0.073
  true spread in the gap sd = 0.102, predicted spread sd = 0.076
  true range 2.19 to 2.72; predicted range 2.15 to 2.47

coal-like input (GR 30, NPHI 0.55, DT 125): training density range 2.04 to 2.79, forest returns 2.15, linear returns 1.60

Assumptions and limitations

  • The relation between the predictors and the target in the gap is the same as in the training samples. This is not tested by validation that mixes depths.
  • The predictors are not damaged in the gap and carry the information needed. If the thing that is missing (a thin dense bed) is invisible to them, it is invisible to the model.
  • The training samples are good. A model trained on damaged samples reproduces the damage.
  • The model can only produce values inside the range it saw. It cannot extrapolate to a rock that is not in the training window.
  • The infilled curve is read as an estimate with an error equal to the blocked-validation error, at best.

QC checks

  • Validation was done on whole depth blocks, and the blocked error is reported next to the random-split error as a check on leakage.
  • The infilled curve joins the measured curve at the edges of the gap with no step, and its spread is compared with the measured spread: a much smoother infill has lost information.
  • Infill is flagged, and quantitative results either mask it or carry a wider uncertainty over it.
  • No infill over a coal, salt or other rock that is absent from the training window.
  • The result is compared with a simple method (linear regression, or just a null). If a forest is no better than a line, use the line.
  • Predictors do not include curves from the same damaged interval or other repaired curves.

Going Deeper

Tree ensembles such as the random forest are popular for infill because they handle non-linear relations and need little tuning. They are also easy to misuse: the accuracy that is reported is a property of the validation scheme as much as of the model, and an infilled curve that looks natural invites being treated as data. The deeper limitation is that a model fitted to the other logs of the same well gives nothing that those logs do not already contain, so the repaired density is a re-expression of the neutron, sonic and gamma ray and does not constitute a measurement of the rock. More elaborate models (neural networks, sequence models) change the form of the prediction but not this limitation. Where a damaged interval matters for the answer, the better remedies are a re-log, a different tool, core or an offset well.

References

  1. Breiman, L., 2001. Random forests. Machine Learning, 45(1), 5–32.
  2. Roberts, D.R. et al., 2017. Cross-validation strategies for data with temporal, spatial, hierarchical, or phylogenetic structure. Ecography, 40(8), 913–929.

Python reference implementation

Python reference implementation

The Python reference implementation is available to registered users with a verified email address. Register or sign in to view it.