Which digits actually go wrong
Character recognition errors are not random. They cluster around glyphs that look alike at low resolution — 1 and 7, 0 and O, 5 and S, 8 and B, 6 and G — and around the marks that carry meaning without being digits at all.
Those marks are where the expensive mistakes live. A decimal point lost to a speck of dust turns 1,234.56 into 123456. A European document that writes 1.234,56 read with US conventions becomes 1.23. A negative rendered as (450.00) becomes positive if the parentheses are dropped. A currency symbol read as a digit prepends a stray character. None of these are exotic; all of them are caught by arithmetic in one line of code.
- Glyph confusions: 1/7, 0/O, 5/S, 8/B, 2/Z, 6/G — worst on low-DPI and dot-matrix print.
- Separators: 1,234.56 versus 1.234,56 versus 1 234,56.
- Negatives: leading minus, trailing minus, parentheses, or a CR suffix.
- Decimals lost to speckle, or invented by a stray mark.
- Currency symbols merged into the number.
- Digits split across a line break in a narrow column.