Loading header...

Home / DPA Insights / Investment Data Operations

Scaling QIS Analytics without scaling headcount

The Build Vs Outsource Decision Facing Risk-Tech Platforms

The short version

For most QIS and risk-tech platforms, getting holdings data is close to a solved problem. The constraint has moved downstream, to decoding each new disclosure format. When the team studied where onboarding time was going, 60% of it was spent decoding, before a single number had been checked. A cycle that took about 75 minutes now takes 30. The analyst’s job was not removed. It was moved: from decoding layout to deciding whether the draft is right. Own the edge. Rent the rest.

Every QIS (Quantitative Investment Strategies) platform needs holdings data. Thousands of funds. Hundreds of formats. An analyst joined the team and spent his first three weeks doing the same thing. Not the same kind of thing but the exact same task, in a slightly different shape, with a different fund family each time. Find the file. Identify the holdings column. Separate value from quantity. Flag the missing attributes. Then start over with the next one. When the team studied where onboarding time was actually going, 60% of it was here in decoding, before a single number had been checked.

For most QIS and risk-tech platforms, getting holdings data is close to a solved problem. Collection has been made routine by years of work on scrapers, pipelines, and standardization. So, when our onboarding time was studied, the finding was a little uncomfortable. The constraint had moved downstream. The question was no longer how the data was obtained. It was how thousands of funds could be covered without an analyst being added for every few hundred.

The bottleneck was decoding, not checking

The analyst time was spent understanding a new disclosure format, not checking whether the data was right. Before a single number was validated, the source had to be found, the holdings file pulled, the value column told apart from the quantity column, and the missing attributes identified. Necessary work. Also, the same work in a slightly different shape, every single time.

So, the team built for that shape. Files are collected from each fund site by a Python scraper. They are standardized into one clean client format by an AI agent. The funds are then checked against our rules by a QC macro, which surfaces the outliers and returns cleaner data. A cycle that took about 75 minutes now takes 30.

It did not work cleanly the first time. A fund family changed their disclosure structure mid-project and the scraper returned nothing for three days. The fix was straightforward. The lesson was automation that cannot detect its own failures is not automation. It is a quiet liability.

ReadoutInternal onboarding study
60% of onboarding time spent decoding, before a single number had been checked
30 min cycle time now, down from about 75 minutes
3 days the scraper returned nothing after a disclosure structure change. Automation that cannot detect its own failures is a quiet liability

The check is the point, not the overhead

The team was careful about one thing. A mapping model is perfectly capable of being confidently wrong. A settlement value can be mapped to market value and handed over as if nothing is off, and the error stays hidden until it surfaces in a risk number of weeks later. So, the analyst’s job was not removed. It was moved. Instead of decoding layout, they now decide whether the draft is right.

Own the edge, rent the rest

That is what makes build versus outsource mostly the wrong question. The useful one is narrower. The scraping logic, data models, validation rules, and analytics methods are the actual edge, and they are kept in-house. Document understanding, orchestration, and first-pass checks can now be bought and adapted rather than built from scratch.

The teams that scale coverage well over the next few years will not be the ones with the biggest operations headcount. They will be the ones honest about which of their work is real expertise and which is repetitive decoding that a machine can draft and a person can check. The first is worth protecting. The second was never a good use of an analyst’s day.

Table 01 / Own the edge vs rent the rest
Question Own the edge Rent the rest
What it covers Scraping logic, data models, validation rules, analytics methods Document understanding, orchestration, first-pass checks
How to source it Kept in-house Bought and adapted rather than built from scratch
What it is Real expertise — worth protecting Repetitive decoding that a machine can draft and a person can check
The analyst’s role Decide whether the draft is right No longer spent decoding layout

Separate the edge from the decoding.

Which of your data work is expertise, and which is decoding?

If your team is weighing build versus outsource for holdings data, we will start by separating the edge you should keep in-house from the repetitive decoding a machine can draft and a person can check.

Book a working session

Frequently asked questions

What is the real bottleneck in scaling holdings data coverage for QIS platforms?

Decoding, not collection or checking. For most QIS and risk-tech platforms, collection has been made routine by years of work on scrapers, pipelines, and standardization. The analyst time goes into understanding each new disclosure format: finding the source, pulling the holdings file, telling the value column apart from the quantity column, and identifying the missing attributes. When the team studied where onboarding time was going, 60% of it was spent decoding, before a single number had been checked.

How can AI reduce the time it takes to onboard fund holdings data?

By building for the repeating shape of the work. Files are collected from each fund site by a Python scraper, standardized into one clean client format by an AI agent, and then checked against the team’s rules by a QC macro that surfaces the outliers and returns cleaner data. A cycle that took about 75 minutes now takes 30.

Should an analyst still review AI-mapped holdings data?

Yes. A mapping model is perfectly capable of being confidently wrong. A settlement value can be mapped to market value and handed over as if nothing is off, and the error stays hidden until it surfaces later in a risk number. So the analyst’s job is not removed; it is moved. Instead of decoding layout, the analyst decides whether the draft is right.

Why does holdings data automation need to detect its own failures?

When a fund family changed its disclosure structure mid-project, the scraper returned nothing for three days. The fix was straightforward. The lesson was that automation that cannot detect its own failures is not automation. It is a quiet liability.

What should a risk-tech platform keep in-house and what can it outsource?

The scraping logic, data models, validation rules, and analytics methods are the actual edge, and they are kept in-house. Document understanding, orchestration, and first-pass checks can now be bought and adapted rather than built from scratch. The useful question is not build versus outsource, but which work is real expertise and which is repetitive decoding that a machine can draft and a person can check.

Abhijit Aher
Abhijit Aher

Manager, Research and Data Operations,

Decimal Point Analytics Pvt Ltd

Abhijit Aher is Manager, Research and Data Operations at Decimal Point Analytics. This article is about where QIS holdings-data time actually goes, and why the useful question is not build versus outsource, but which work is real expertise.