Get our report on investing trends!
By providing your email, you will shortly receive the latest report from Pepper.
The throughput ceiling of manual GP report processing is a structural constraint on secondaries portfolio intelligence at scale — and AI-powered format-independent data ingestion is the only architectural response that eliminates it rather than managing it.
The secondaries portfolio management data infrastructure problem is simpler to state than most technology problems: GP reports arrive from dozens of external sources in dozens of different formats, and the portfolio intelligence the manager needs to monitor risk, price transactions, and report to LPs requires all of that data to be normalised into a single, consistent, queryable dataset. The manual version of this process has a throughput ceiling. The AI-powered version does not.
This paper describes the cost of the manual baseline, the architecture of the AI-powered alternative, the quality control infrastructure that makes it trustworthy, and what portfolio intelligence becomes possible when the ingestion problem is solved.
A skilled operations analyst processing GP quarterly reports handles approximately five to eight reports per day, depending on report complexity, format variety, and depth of data extraction required. At this rate:
5–8 GP reports processable per analyst per day through manual extraction
100+ reports an AI format-independent ingestion system processes simultaneously as they arrive
The throughput ceiling is not primarily a cost problem. It is a data currency problem. A process that cannot keep up with report volume produces a portfolio view that is always partially stale — and the positions with the stalest data are, systematically, the positions where current data matters most.
A secondaries data ingestion system that requires format-specific mapping for each GP reporting template must be reconfigured every time a GP changes their report format and set up from scratch for every new GP added to the portfolio. This creates maintenance overhead that grows with the portfolio. A system with generalised extraction capability — one that identifies NAV, IRR, MOIC, DPI, TVPI, capital account balance, unfunded commitment, and underlying company list regardless of how they are labelled, positioned, or structured in the source document — scales without proportional maintenance overhead. The technical investment in generalised extraction is significantly greater than in format-specific mapping; the long-term operational cost is significantly lower.
AI extraction is accurate but not infallible. Extractions that the AI is uncertain about — unusual formatting, ambiguous metric labels, values that fall outside expected ranges for a fund of this type and vintage — should be flagged for human review rather than accepted automatically. The exception rate in a well-calibrated system is low (typically under 5% of extracted data points), but the consequences of passing incorrect data through without review can be significant. Best-in-class systems flag exceptions, route them to a human reviewer through a structured workflow, and track resolution for audit purposes.
Every extracted metric should be compared automatically against the prior-period value for the same fund. Large period-over-period movements — NAV changes that are inconsistent with reported portfolio performance, distributions that do not match expected capital activity, MOIC changes that imply implausible returns for the period — should be flagged for analyst verification before they enter the portfolio data model. This automated cross-validation catches extraction errors that the primary AI might miss and GP reporting anomalies that are important signals for portfolio management.
Once GP report data is normalised into a unified, structured, continuously updated dataset, three categories of secondaries portfolio intelligence become achievable that are not feasible through manual processing.
A secondaries portfolio with 100 underlying fund positions may hold exposure to the same underlying company through five or six different fund positions — positions in funds managed by different GPs, with different entry vintages, representing different portions of the same company’s capital structure. This concentration is invisible when each GP report is processed individually. It becomes visible — and manageable — when all underlying portfolio company data is aggregated into a single queryable dataset. For a secondaries manager managing LP reporting to institutional LPs who want underlying company exposure breakdowns, this capability is a reporting requirement, not just an analytical one.
Disaggregating secondaries portfolio returns by the vintage year of each underlying fund position, the underlying strategy (buyout, growth, credit, infrastructure), and the quality tier of each GP requires consistent, comparable data across all positions. Manual processing — which produces inconsistently structured data from heterogeneous sources — makes this analysis imprecise at best and impossible at institutional portfolio scale. AI-normalised data makes it straightforward.
Systematically tracking GP-reported NAVs over time and comparing them to comparable funds and underlying portfolio company performance requires a database of normalised GP-reported data spanning multiple reporting periods and multiple fund positions. This database exists, in a consistent and queryable form, only if the data ingestion process has been systematic and the data model has been consistently structured. On that foundation, AI can identify mark quality signals — systematic mark stability through volatile periods, persistent deviation from comparable fund marks, mark movements inconsistent with underlying portfolio performance — that are invisible to manual analysis at the scale and frequency needed to be actionable.
Pepper ingests GP reports in any format and extracts consistent, structured data automatically, with exception flagging and automated cross-validation. Cross-portfolio concentration analysis, performance attribution, and NAV quality monitoring all operate on the normalised, unified data that ingestion produces. The data ingestion architecture is the foundation on which secondaries portfolio intelligence is built — and it was designed specifically for the scale and format heterogeneity of the secondaries use case.
Sign up for our newsletter to receive biweekly updates on the world of asset management, delivered straight to your inbox.
By providing your email, you will shortly receive the latest report from Pepper.
In a 45-minute session, we'll walk you through how Pepper handles the workflows your team runs today — deal management, portfolio monitoring, fund operations, or LP reporting. You pick the priority.
Not a sales call. A 30-minute conversation with a Pepper practitioner about where your operation is today, where the pressure points are, and whether a platform approach makes sense for your stage of growth.