Pepper — private credit investment platform Pepper
Secondaries Thought Leadership

The Secondaries Data Problem: Why Portfolio Visibility Breaks Down at Scale — and How to Fix It

The larger a secondaries portfolio becomes, the less the manager actually sees — because portfolio visibility in secondaries is constrained not by the manager's own data management, but by the quality and timeliness of GP reporting from dozens of external sources.

 

Here is the paradox at the centre of secondaries portfolio management: portfolio visibility decreases as portfolio size increases. Not proportionally — faster than the growth. At 20 underlying fund positions, the secondaries manager has reasonably current data on most of the portfolio. At 100 underlying fund positions, the manager has current data on a fraction of it — and the fraction with stale data is not random. It is concentrated in the positions with the least sophisticated GPs, which are also the positions where current data matters most.

This is not a description of inadequate secondaries portfolio management. It is a description of what happens when portfolio size outpaces the capacity of manual GP report processing — which is what happens in every secondaries operation that has grown past the $1 billion AUM mark without a fundamental change in data infrastructure.

Why secondaries has a data problem that no other private markets strategy has

In direct lending, private equity, or mezzanine, the manager controls the data about their investments. The credit agreements, the borrower financial statements, the covenant calculations, the portfolio company financial data — it exists in the manager’s own systems. The data challenge is operational: how to process and organise it efficiently.

In secondaries, the manager does not control most of the underlying data. They own interests in funds managed by dozens of different GPs, each with their own reporting timelines, formats, disclosure standards, and data quality levels. The underlying portfolio companies — potentially thousands of them, in a large fund-of-funds secondaries portfolio — report to GPs. The GPs report to the secondaries manager. The secondaries manager’s view of their portfolio is mediated by two layers of intermediation, and the quality of that view is constrained by the weakest link in the GP reporting chain.

The weakest link is not the average GP. The weakest link is the least sophisticated GP in the portfolio — the emerging manager or the mid-market fund with limited fund administration resources — and that weakest link typically manages the assets with the most variable quality. The positions where the secondaries manager most needs current, accurate data are the positions most likely to have stale, incomplete, or late GP reporting.

How the visibility gap compounds at each portfolio size threshold

At 10–20 underlying fund positions: Manual GP report processing is manageable. Each quarterly report gets processed by an analyst within two weeks of receipt. The portfolio view is reasonably current for most positions. Data currency problems are exceptions, not the norm.

At 50 underlying fund positions: The manual processing bottleneck creates systematic data currency problems. Not every GP delivers on the same timeline. The analyst processing 50 quarterly reports — each in a different format, from a different GP portal, with a different disclosure scope — is spending the majority of their time on data collection and entry, with limited capacity for analysis. Some positions are current. Others are one reporting cycle behind. A few are two cycles behind.

At 100+ underlying fund positions: The positions with the most current data are the positions with the most sophisticated GPs — who are also, typically, the positions where the underlying portfolio quality is highest and the monitoring need is lowest. The positions with the stalest data are the positions with the least sophisticated GPs — where the underlying portfolio quality is most variable and the monitoring need is highest. The data problem is correlated with risk.

Quote Icon

The secondaries data problem is not that GPs report badly. It is that the volume and heterogeneity of GP reporting, at institutional portfolio scale, exceeds the capacity of any manual processing system to maintain current, consistent data across the full portfolio. The structural fix is format-independent AI data ingestion.

What AI-powered format-independent data ingestion changes

The structural fix for the secondaries data problem is AI-powered document ingestion: the ability to process GP reports in any format — PDF quarterly reports, Excel data packages, GP portal exports, ILPA-formatted data packages, CSV extracts — and extract consistent, structured data automatically, without requiring a manual mapping step for each new GP reporting format.

The architecture of this ingestion layer determines whether it scales with the portfolio or creates new bottlenecks. A system that requires format-specific mapping for each GP reporting template must be reconfigured every time a GP changes their report format, and it must be set up from scratch for every new GP position added to the portfolio. A system with generalised extraction capability — one that identifies NAV, IRR, MOIC, DPI, TVPI, capital account balance, unfunded commitment, and underlying company list regardless of how they are labelled or positioned in the source document — scales with the portfolio without proportional maintenance overhead.

Quality control infrastructure is as important as extraction capability in a secondaries data ingestion system. AI extraction is accurate but not perfect. Best-in-class secondaries portfolio management platforms include exception flagging for extractions the AI is uncertain about, automated cross-validation of extracted metrics against prior-period data, and human review workflows for flagged exceptions. The AI handles the volume. Human review handles the exceptions. Both are necessary for a secondaries data infrastructure that a portfolio manager can trust.

A note on Pepper’s approach

Pepper ingests GP reports in any format — PDF, Excel, CSV, ILPA-formatted data packages — and extracts consistent, structured data automatically, with exception flagging for cases where the AI flags uncertainty. Cross-portfolio concentration analysis, performance attribution by vintage and GP quality tier, and NAV quality monitoring all operate on the normalised, unified data that this ingestion produces. The secondaries data problem — GP report aggregation at institutional portfolio scale — is the specific operational challenge the Pepper data ingestion architecture was built to address.

Related articles

Aren't you just a little curious?

Sign up for our newsletter to receive biweekly updates on the world of asset management, delivered straight to your inbox.