untethered atom · EBSD & TKD

EBSD & TKD · Part 5 of 8

Why does EBSD cleanup change my grain count?

Cleanup is where fake grains go to die (and sometimes real ones, oops).

Every published EBSD map has been cleaned. The raw file has holes where indexing failed and speckle where the vote elected the wrong crystal, and there are standard algorithms that make both vanish. The algorithms are fine. What deserves scrutiny is that every one of them works by replacing measurements with guesses. This page has something a real experiment never has: the true answer, so you can watch exactly when the guessing turns into invention.

Cleanup does not recover data. It manufactures agreement with its own assumptions, which is fine, right up until it isn't.

01

What raw data looks like

Here is a microstructure whose every pixel is known exactly (parents and twin lamellae, the same cast as part 3), and next to it, what a real scan of it would return. Two corruptions do almost all the damage in real data, and you control both.

Non-indexed points cluster at grain boundaries, where the interaction volume straddles two crystals and the pattern is two patterns superimposed, plus scratches, contamination and dust scattered anywhere. The software returns nothing, honestly. Mis-indexed speckle is subtler: the triplet vote from part 1 elects a wrong-but-confident orientation (pseudosymmetry, overlap, bad luck), sprinkling single pixels of a foreign orientation through otherwise clean grains. Both arrive exactly where the confidence index already warned you they would.

Truth vs measurement The truth is what an experiment never has
Non-indexed Speckle = single pixels in a foreign colour
Indexing rate
% of pixels with an answer
Wrong answers
% of indexed pixels · invisible on the map
Flip between the views. The holes announce themselves; the speckle you can squint past; and the wrong answers that matter most, confidently mis-indexed pixels, look exactly like data. Nothing on the measured map distinguishes them except, sometimes, a low CI.
02

The cleanup machine

Vendors ship a stack of cleanup passes (grain dilation, grain CI standardisation, neighbour-orientation correlation), and they compose. The machine below runs the same logic in its plainest form: first fill each hole with the majority of its neighbours, then keep going: reassigning any pixel that disagrees with its neighbourhood. One slider, from "untouched" to "thoroughly laundered".

One slider, from raw to laundered Watch the counter, not just the picture
Pixels rewritten
% of the map the software changed
Of those, guessed right
% agreeing with the truth
Levels 1–3 mostly fill holes, and the guesses are mostly right: a hole between two pixels of grain A probably was grain A. From level 4 the machine starts overruling measured pixels, and by the top of the slider it is straightening boundaries and eating visibly into the finest twin lamella. The map gets prettier the entire time.
Go deeper: the real algorithms, by name

Grain dilation grows indexed grains into non-indexed space until the holes are gone: the fill stage above. Grain CI standardisation assigns every pixel in a grain the grain's best CI, so that a subsequent CI filter throws away isolated low-confidence points but keeps grains; it rewrites no orientations itself, but it decides what the next pass may rewrite. Neighbour orientation correlation reassigns pixels that disagree with most of their neighbours: the smoothing stage above, and the one that eats real features. Newer pattern-space methods (NPAR and its cousins) average raw patterns before indexing instead of editing orientations after: better behaved, same obligation to say so.

Every package logs none of this in the figure. The map that emerges carries no watermark saying "14% of these pixels are interpolation". That watermark is the caption's job, which is the whole point of this page.

03

What it does to your statistics

The picture is not the deliverable; the numbers are. Because this page knows the truth, it can plot what no experiment can: the actual error of every statistic, at every cleanup level. The curves below re-run part 3's grain count and twin fraction on the cleaned map, at every level of the slider, and compare against the known answer.

The curves nobody can plot for real data Dashed lines are the truth
Grain count Σ3 twin length fraction Pixels wrong vs truth True values (dashed)
Best level for this dataset
minimum pixel error, knowable only with truth
Grain count at current level
true:
Σ3 fraction at current level
true:
The pixel-error curve is U-shaped: cleanup genuinely helps, then genuinely harms, and the turn happens while the map is still improving to the eye. In a real experiment this curve is invisible. You are choosing a point on it blind. That asymmetry is the entire ethics of cleanup.
The honest protocol

State every cleanup step and its parameters in the methods. Report the raw indexing rate. Keep a raw map in the supplement. Compute boundary and grain statistics only from measured pixels where feasible, and check that the conclusion survives at zero cleanup and at double your chosen level. If a result appears only after cleaning, it is a property of the cleaning.

A trap worth naming

Cleanup is validated by eye against expectations, which makes it a machine for confirming expectations. The nastiest failure mode is not noise surviving; it is a real, unexpected feature (a fine twin, a recrystallised nucleus, a second phase) being scrubbed because the algorithm, tuned until the map "looked right", counted it as noise. Watch the finest lamella in the demo above being eaten alive over the last few slider levels. Yours will be too.

Sources & further reading

Cite this page: Tripathy, Manisha. “Cleaning data honestly.” untethered atom, 2026, https://untetheredatom.com/ebsd/ebsd-5-cleaning-data-honestly.
BibTeX
@misc{tripathy2026cleaningdatahonestly,
  author = {Tripathy, Manisha},
  title  = {Cleaning data honestly},
  year   = {2026},
  howpublished = {\url{https://untetheredatom.com/ebsd/ebsd-5-cleaning-data-honestly}},
  note   = {Interactive teaching resource}
}
Last updated 12 August 2026.