Regional geochemistry · Argentine Puna

Does the zone hold copper?

Two public sources covered the ground: about 3,500 fifty-year-old readings in coarse steps, and precise modern re-analyses of a third of them. We calibrated one against the other, batch by batch, and mapped copper around the project with a margin of error that held when tested.

Red-orange cut faces and haul roads of a mine seen from directly above

The question

Does the zone hold copper? A client with ground in the Argentine Puna wanted to know what the public record already said about copper in and around it, and how far that record could be trusted, before paying for field work.

The data available

We searched the public record and found two sources from the national geological survey. Both describe the same stream-sediment samples, collected in a regional survey in the 1960s and 1970s.

Original resultsRe-analysis
SamplesAbout 3,500, a few dozen of them within 5 km of the projectAbout one third of them, and about one third of those near the project
WhenAt the time of the surveySome thirty years later, on archived splits
MethodMostly visual colorimetry after a partial extractionICP emission spectrometry after a multi-acid digestion
PrecisionSteps of 5 ppm0.1 ppm
What it capturesPart of the metal in the sampleNearly all of it

Neither source answers the question alone. The precise one is too sparse around the project, and the dense one is too coarse. The work was to find out whether the few precise values could be used to improve the estimate from the many coarse ones.

Three more public sources supported the analysis: a 90 m digital elevation model, a 1:250,000 geomorphological map with its drainage, and about 7,000 samples measured both ways on eight neighbouring map sheets.

What was wrong with the data, and what was done about it

ProblemHow it was found and handled
Sample positions were wrongStream sediments should lie on streams. With the published coordinates they were no closer to the stream lines of the elevation model than random points. A rigid shift was fitted for each survey batch against the drainage network: about 1.7 km overall, in the same direction in all 22 batches large enough to fit. The median distance to a stream fell from 365 m to 121 m, and the upstream catchments then matched the sampling density stated in the survey report. Each sample keeps a digitising error of 100 to 300 m, so the maps were redrawn with an alternative correction to see what moves: contours shift by a few hundred metres and keep their shape.
Coarse readingsAbout 95 % of the original values are multiples of 5 ppm, and five values hold most of the samples. Each reading was treated as an interval, not as a number (an interval-censored likelihood).
Values below detectionRecorded as zero, under two notations that change from batch to batch. They were modelled as left-censored data in two ways. Both made the predictions of hidden values worse, so these readings were left out of the copper map.
The two methods do not measure the same thingA partial extraction recovers less metal than a near-total digestion, so no 1 : 1 relation was assumed. Linear, ratio, offset and power-law relations were compared by the Akaike information criterion and by five-fold cross-validated likelihood. A power law won.
Agreement between operators and batchesThe survey was run in 27 batches, each with its own detection limit, reporting habits and offset; visual colorimetry depends on who reads the colour. The relation between old and modern values was therefore fitted batch by batch in one hierarchical model. Rock type was tested as an alternative explanation, at the sample site and over the whole upstream catchment, and did not explain the differences; the batch did.
No duplicates to measure repeatabilitySubstitute duplicates were built from the drainage network: 639 pairs of neighbouring samples on the same stream. Old readings of such neighbours agree to within one 5 ppm step in 92 to 100 % of cases at background levels, so the old method is repeatable, with a noise of 2 to 3 ppm per reading. Its disagreement with the modern method is a matter of scale, not of carelessness.
The re-analysed third was not chosen at randomThe survey selected it by exploration interest, using the old values. We tested the selection: it did not favour high values, it varied from 17 to 45 % between batches, and it left out most of the extremes. Regional percentiles were computed with cell declustering, and old readings higher than any that were re-analysed in their batch received their own rule and their own test (step 9 below).
Few doubly-measured samples in some batchesThe calibration borrowed strength from the eight neighbouring sheets through partial pooling: each batch is estimated from its own pairs and, where these are few, from the population of 129 batches.
Discordant pairsSome pairs disagree far beyond any analytical scatter, and one such pair can flatten the fitted relation of its whole batch (a leverage point). Pairs more than four robust standard deviations from the others with the same reading were kept out of the fit and listed for review.
Mixed sample materialsThe tables include alluvium and soil beside stream sediment, without explanation. Alluvium proved comparable and was kept with a wider error. Soil read about 30 % lower and was kept out of the map.
Different splits, thirty years apartThe two values of a sample come from different splits, and the modern one from 0.1 to 0.2 g of material. The resulting scatter was estimated twice, from the substitute duplicates and from the nugget of the variogram, and carried into the estimate as measurement error.

What was done

  1. Audit. Every count in the data description was reproduced independently before any analysis: 53 checks, all passed.
  2. Terrain analysis. The elevation model was depression-filled, with real closed basins kept as sinks, and flow was routed across it. This gave the corrected positions, the upstream catchment of every sample, the rocks inside each catchment and the order of samples along each stream.
  3. Pair analysis. On the samples measured both ways, the rank correlation (Kendall tau-b) between old and modern values is 0.47 for copper. The modern value is the old value plus a few ppm in every batch, with an offset that changes from batch to batch.
  4. Variography. Empirical variograms and a nested covariance model fitted by restricted maximum likelihood gave the spatial structure of the modern values. Once the batch is accounted for, the copper discrepancy between the two methods is pure nugget: the conversion errors of neighbouring samples are independent, which is what the next two steps require.
  5. Hierarchical calibration. An interval-censored power-law regression of the old reading on the modern value, with random intercepts and slopes for every batch and every sheet, fitted to the doubly-measured samples of all nine sheets: about 7,100 after screening, in 129 batches. Inverting it turns each old reading into a value on the modern scale with its own error variance.
  6. Data fusion. Kriging with variance of measurement error: a modern value counts almost in full, a converted old reading counts according to its own margin, and every place is estimated from its own value and its neighbours. A converted reading cannot overrule a real measurement next to it.
  7. Benchmarks. The same estimate was made in four other ways: modern values alone, one rescaling rule for the whole sheet (quantile mapping), the converted value without its neighbours, and cokriging with a linear model of coregionalisation.
  8. Cross-validation. The modern value of some samples was hidden, the calibration was refitted without them, and each method had to predict them. Three designs: random folds within every batch, spatial blocks of 10 km, and whole batches left out. Predictions were scored by the continuous ranked probability score, mean absolute error, the area under the ROC curve for telling a high sample from an ordinary one, and the coverage of the 90 % intervals, with bootstrap intervals on the differences.
  9. Extrapolation test. Around the project some old readings are higher than any that were re-analysed in their batch, so the calibration is unchecked there. Samples with the highest readings of each batch were hidden and predicted under several rules, and the rule that tested best was applied.
  10. Maps. For copper: the estimate on a 500 m grid, the estimate relative to the background of the rocks upstream, the margin of error, the probability of exceeding the regional 90th percentile, and the catchments coloured by their sample. Ground farther than 4 km from any sample is left blank.

The result

For copper, the precise few do improve the coarse many, and the improvement was measured against hidden values, not assumed.

CopperModern values aloneBoth sources combined
Error at a sample that has only an old readingBaseline13 % smaller; 19 % smaller within 20 km of the project; 30 to 38 % smaller where the old reading is 20 ppm or more
Telling a high sample from an ordinary oneRight 76 times in 100Right 86 times in 100
Coverage of the project's zoneGaps: two thirds of the samples within 5 km have only an old readingGaps filled from the old readings nearby
Stated 90 % rangeContained the hidden value 90 to 92 times in 100

With that estimate, the answer to the question is a qualified yes. Copper is above the district's typical level over most of the area around the project. The project ground maps in the upper quarter of the district, at 1.3 to 1.6 times what its rocks normally give. A zone in the top tenth lies a few kilometres away, and the highest modern values within 20 km are five to six times the rock background. These are signs of copper enrichment in stream sediment: high for the district, and low in absolute terms.

The same tests showed that the obvious shortcut does harm: one rescaling rule for the whole sheet gave errors 30 to 125 % larger than ignoring the old values altogether.

Have data and a decision to make?

Tell us what you have and what you need to decide. We will say what your data can support, and what it cannot, before any work starts.

Discuss your project