Loading header...

Home / DPA Insights / Financial Reference Data

Grammatically Fine, Financially Wrong

A fund name, a fee, a share class identifier is only useful if it is correct every time, not most of the time.

The Short Version

The target was 15,721 Korean mutual fund share classes. The first validation pass sat at 70 to 80 percent accuracy — a wall for financial reference data. The real problem was not extraction. It was translation that read correctly but wasn’t. It was solved by a loop, not one better prompt. Every data point is checked by a human against the original document, permanently. Validated accuracy reached 99 percent. Multilingual financial data work fails less often at the language layer and more often at the definitions layer.

Fifteen thousand seven hundred and twenty-one Korean mutual fund share classes. That was the target when the project started, and for the first few weeks, it looked achievable in a straightforward way: find the prospectuses, extract the data, translate it into English. Then the team ran their first validation pass, and the accuracy sat at 70 to 80 percent.

The real problem was not extraction

For most automation projects, that number would be a milestone. For financial reference data, it was a wall. A fund name, a fee, a share class identifier is only useful if it is correct every time, not most of the time. The team’s real problem was not extraction. It was translation that read correctly but wasn’t. Korean entity names, for example, would come back as literal English phrases instead of the legal or commercial name a fund actually uses. Grammatically fine, financially wrong.

That distinction changed the question the team was asking. Instead of “can the system pull this information out of the document,” they started asking whether it could pull the correct information, consistently, in a form usable as structured financial data. Those are different engineering problems, and the second one does not get solved by writing one better prompt.

Building a loop: extract, validate, investigate, refine

It got solved by building a loop: extract, validate against the source, investigate why a field failed, refine the prompt or the logic, extract again. Every data point was still checked by a human against the original document, not as a stopgap while the automation matured, but as a permanent part of the system, because it was the only way to tell whether a wrong answer was an extraction problem, a translation problem, or a definitions problem specific to Korean fund structures. The reviewer was not checking whether the machine had read the document. They were checking whether it had understood it. Those are not the same question.

A repeatable method

Validated accuracy climbed from that first 70 to 80 percent to 99 percent. And the human check, the part you might expect to be the bottleneck, got faster as the loop matured, with review time per fund falling from 12 minutes to 3. What the team really built was a repeatable method, drawn from source discovery, field-specific prompting, iterative validation, and human review, and it has since been reused on two more countries.

ReadoutKorean mutual fund share classes
15,721 Korean mutual fund share classes in the original target
99% validated accuracy, up from 70 to 80 percent on the first pass
3 min review time per fund, down from 12 minutes as the loop matured

The definitions layer, not the language layer

The lesson generalizes past Korea. Multilingual financial data work fails less often at the language layer and more often at the definitions layer: what a field means, in that market’s convention, in that document’s format. Source discovery and human validation are not the parts of the workflow you automate right away once the model gets good enough. They are the parts that tell you whether the model is actually right.

Table 01 / Reading the document vs understanding it
Question Reading the document Understanding the document
The question Can the system pull this information out of the document? Can it pull the correct information, consistently, in a form usable as structured financial data?
A Korean entity name Comes back as a literal English phrase The legal or commercial name a fund actually uses
Where it fails The language layer The definitions layer: what a field means, in that market’s convention, in that document’s format
How it gets solved Writing one better prompt (it doesn’t) A loop of extraction, validation against the source, investigation and refinement, with human review

Find where the wrong answers come from.

Is your multilingual data read, or understood?

If your team is building reference data from documents in other languages, we will start by finding where your wrong answers come from — extraction, translation, or definitions — before anyone talks about tooling.

Book a working session

Frequently asked questions

Why is 70 to 80 percent accuracy not good enough for financial reference data?

A fund name, a fee, a share class identifier is only useful if it is correct every time, not most of the time. For most automation projects, 70 to 80 percent accuracy would be a milestone. For financial reference data, it is a wall.

Why does machine translation fail on Korean mutual fund data?

Because translation can read correctly without being correct. Korean entity names, for example, came back as literal English phrases instead of the legal or commercial name a fund actually uses. Grammatically fine, financially wrong.

How do you improve accuracy in multilingual fund data extraction?

Not by writing one better prompt. The team built a loop: extract, validate against the source, investigate why a field failed, refine the prompt or the logic, and extract again. Validated accuracy climbed from 70 to 80 percent on the first validation pass to 99 percent.

Is human review still needed once AI extraction becomes accurate?

Yes. Every data point was checked by a human against the original document, as a permanent part of the system rather than a stopgap. It was the only way to tell whether a wrong answer was an extraction problem, a translation problem, or a definitions problem specific to Korean fund structures. The check also got faster as the loop matured, with review time per fund falling from 12 minutes to 3.

Does the method work for markets beyond Korea?

Yes. The repeatable method, drawn from source discovery, field-specific prompting, iterative validation, and human review, has since been reused on two more countries. Multilingual financial data work fails less often at the language layer and more often at the definitions layer: what a field means, in that market’s convention, in that document’s format.

Francis DSouza
Francis DSouza

Senior Research Analyst and Team Leader,

Decimal Point Analytics Pvt Ltd

Francis DSouza is Senior Research Analyst and Team Leader at Decimal Point Analytics. This article is about what translating Korean mutual fund reference data taught the team: the work fails less often at the language layer, and more often at the definitions layer.