CamPetro

Lithofacies Modelling

On this page

Purpose

Lithofacies modelling gives each depth sample a rock type from the logs: shale, sandstone, limestone, dolomite, or a finer scheme. A facies curve has several uses. It selects the interval in which a porosity, saturation or permeability model applies, it controls the matrix parameters and the cutoffs, and it is the main input to a geological model of the reservoir. A good facies curve reflects the geology and also the data: it is only as reliable as the logs, the normalization and the labels behind it. This step describes three families of method, rules, clustering and supervised classification, and the choice of curves and the handling of missing data that all three share.

Position in the workflow

Upstream. Facies depend entirely on the quality of the logs. Normalized, repaired logs from Stage 1 are the first requirement, and the facies track is null wherever a needed curve is null. Clay volume and porosity are not required, but a facies curve is often computed before them so that the matrix and the volume models can follow the facies.

Downstream. The facies curve feeds the choice of matrix density and clay parameters in the porosity and mineral models, the choice of \(m\), \(n\) and the saturation model in Water Saturation, the permeability transform, and the zone-by-zone cutoffs. It also supplies the labels for the geological model and for well-to-well correlation.

Error propagation. A misclassified sample is given the parameters of the wrong rock. A limestone called a sandstone gets a density porosity from a matrix 0.06 g/cm³ lighter, which adds about 3 porosity units, and a saturation computed with the exponents of a sand. The errors are largest at facies boundaries and for rare facies, which are the ones with the fewest samples to check them against. In the worked example on the deterministic page, the overall accuracy of a good rule table is 0.84, but its limestone recall is 0.59. Report the confusion matrix, not one number.

Key concepts

A facies is not a cluster. A Facies code is a number. It becomes a facies only when it is tied to a named rock type by a rule, a core description or a geologist who looked at the centre of the cluster.

Three families of method. Deterministic rules are thresholds on crossplots, written as an ordered table. Unsupervised clustering groups samples by similarity with no labels, with k-means clustering as the standard algorithm. Supervised classification learns from core-described facies and is validated by held-out blocks.

Neutron-density separation. Neutron-density separation is the log-derived quantity that separates shale, sandstone and carbonate on a single axis, and the backbone of most rule tables.

Validation without leakage. Neighbouring samples are not independent. Any test must hold out whole depth intervals or whole wells: see Depth-block cross-validation.

Class imbalance and null propagation. Rare facies are missed unless the evaluation looks at them (Class imbalance), and missing inputs must give missing outputs (Null propagation). See Curve Selection and Missing Data.

Method selection guide

Method Inputs Use when Strengths Weaknesses
Deterministic Rules from Crossplots GR, density, neutron, optionally resistivity and PEF; thresholds Few lithologies, good logs, a need for transparent and repeatable results No training data, readable, easy to adjust Hand-set thresholds, hard boundaries, overlapping facies are confused
Unsupervised Clustering Standardised curves, number of clusters No core description; exploring a well; finding groups a table missed No labels needed; finds structure in many curves Clusters need naming; the best k is ambiguous; ignores depth; unstable between wells without normalization
Supervised Classification Normalized curves and core-described facies Core description exists and the facies are geological Learns the real facies scheme; gives skill by class Needs labels; leakage inflates skill; rare facies missed; limited to the training conditions
Curve Selection and Missing Data All candidate curves of every well Always, before any method, and again whenever the curve set changes Removes redundancy, makes wells comparable, makes gaps explicit Cannot create information that the logs do not hold

Decision guidance

  • If the lithologies are few and the logs are good, start with a rule table. It is a baseline that every other method has to beat.
  • If there is no core description, cluster with a modest k, name the clusters from their centres, and compare the result with the rule table.
  • If there is core description of several wells, train a classifier and score it with held-out depth blocks or wells, and show the confusion matrix with the recall of each facies.
  • If the facies differ only by texture, grain size or sedimentary structure, the conventional logs cannot separate them. Use image logs or core, or accept a coarser scheme.
  • If curves are missing in some wells, prefer a smaller model validated for that curve set over a filled curve, and return null where neither is available.

Shared parameter picking

Curve set. The same list of curves is used by clustering and by classification, and is chosen once for the project: see Curve Selection and Missing Data.

Standardisation. Mean and standard deviation of each curve from the key well, applied to every well after normalization. Resistivity as log10.

Facies scheme. One list of names and codes for the project, with the same colour in every plot. State which codes are lithologies (shale, sandstone, limestone, dolomite) and which are special (evaporite, coal, casing, organic shale).

Depth filter. A majority filter with a width near the thinnest bed of interest, applied in the same way to every method so that results can be compared.

Required curves. The curves without which the facies track is null, shared by all methods.

Key well and core. The well that provides the thresholds, the cluster centres or the labels, and the core intervals reserved for testing.

Absent other information, a careful generalist would:

  1. Normalize the logs, choose the curve set (gamma ray, density, neutron, log resistivity, and sonic or photoelectric factor if good), drop redundant curves, and decide which curves are required.
  2. Build a rule table on the key well from crossplots and compare it with core. This is the baseline.
  3. Cluster the standardised curves for k from 2 to about 8 with several seeds, pick k with the silhouette and the geology, name the clusters from their centres, and compare them with the rule table.
  4. Where core description exists, train a supervised classifier and report block or leave-one-well-out accuracy, with the confusion matrix and the recall of every facies.
  5. Choose the method that is best against core, or the rule table if none beats it clearly. Smooth with a majority filter and return null where required curves are null.
  6. Check the facies proportions and the facies track in every well before the curve is used to drive matrix and saturation parameters.

Combining methods

The methods are not rivals. A rule table is a way to name clusters, since the same thresholds that define a facies by rule can label a cluster from its centre. A supervised model can be trained on core and then compared with the rule table to find the intervals where they disagree, and those are the intervals to look at in the core. A cluster that matches no rule is either a new facies or a log problem, and finding out which is often the most useful result of the exercise. When two methods disagree, the disagreement is a map of uncertainty. Combining by vote is possible, but the methods share their inputs and fail together when a log is bad, so the vote is less independent than it looks. Prefer to report one method, with the others as checks.

QC of results

A good result:

  • has a clear name for every facies code, tied to core or to a rule,
  • compares with core by a confusion matrix on held-out depth blocks or wells, with recall by class,
  • gives facies proportions that the geology supports in every well,
  • is null exactly where the required curves are null,
  • changes in thick, geologically plausible intervals, not every few samples, and
  • is stable under a different seed, a different filter width and the removal of a redundant curve.

Signs of a bad result: shale in clean sandstone with a clean gamma ray, limestone in a washout, a facies that appears only in one well, a skill figure from a random split, and a smooth, complete track in an interval where a curve is missing.

Common pitfalls

  • Reading a cluster number as a lithology without naming it from the centre or the core.
  • Clustering curves that have not been standardised, so that one curve dominates.
  • Choosing k by the silhouette or the elbow alone.
  • Testing a classifier with a random split of samples, so that neighbouring samples of one bed are on both sides.
  • Quoting overall accuracy when a rare facies has a recall near zero.
  • Fitting the standardisation or the thresholds on the whole well, including the test depths.
  • Applying a model to a well with another tool or another vintage without normalization.
  • Filling a missing curve and treating the result as if all curves were present.
  • Keeping two curves that are one curve in two forms, such as bulk density and density porosity.
  • Letting a log artefact, such as a washout or casing, form a facies of its own.

Going Deeper

Automated lithology from logs goes back to crossplot interpretation, the matrix identification charts of the service companies, and the multi-mineral solutions of the 1970s and 1980s. Clustering and classification were added with computers under the name of electrofacies, and machine learning has made classifiers on many curves routine. The state of practice is that the method matters less than the data: careful normalization, a sensible facies scheme, honest validation and a refusal to answer without the logs improve the result more than a more flexible model. The open problems are shared by all methods. Facies defined by texture, grain size or bedding are poorly resolved by conventional logs. Boundaries are gradational in the rock and sharp in the track. And a facies model trained in one basin rarely transfers to another. Probabilistic outputs, which give a probability per facies and a flag where the probabilities are close, are the most practical improvement on a single label.

Methods in this step