Pepper — private credit investment platform Pepper
Secondaries White Paper

From GP Reports to Portfolio Intelligence: Building a Scalable Data Infrastructure for Secondaries Managers

The throughput ceiling of manual GP report processing is a structural constraint on secondaries portfolio intelligence at scale — and AI-powered format-independent data ingestion is the only architectural response that eliminates it rather than managing it.

 

The secondaries portfolio management data infrastructure problem is simpler to state than most technology problems: GP reports arrive from dozens of external sources in dozens of different formats, and the portfolio intelligence the manager needs to monitor risk, price transactions, and report to LPs requires all of that data to be normalised into a single, consistent, queryable dataset. The manual version of this process has a throughput ceiling. The AI-powered version does not.

This paper describes the cost of the manual baseline, the architecture of the AI-powered alternative, the quality control infrastructure that makes it trustworthy, and what portfolio intelligence becomes possible when the ingestion problem is solved.

The throughput ceiling of manual GP report processing — quantified

A skilled operations analyst processing GP quarterly reports handles approximately five to eight reports per day, depending on report complexity, format variety, and depth of data extraction required. At this rate:

  1. A secondaries portfolio with 30 underlying fund positions requires 30 quarterly reports processed — four to six analyst-days per quarter, every quarter, just to maintain current data.
  2. A 60-position portfolio requires eight to twelve analyst-days per quarter of processing time.
  3. A 100-position portfolio consumes more than two full analyst-weeks per quarter — before a single hour of analysis has been performed.
  4. At 150 positions — the scale of meaningful institutional secondaries portfolios — manual processing cannot sustain current data across the full portfolio without a dedicated data operations team that grows proportionally with AUM.

5–8 GP reports processable per analyst per day through manual extraction
100+ reports an AI format-independent ingestion system processes simultaneously as they arrive

The throughput ceiling is not primarily a cost problem. It is a data currency problem. A process that cannot keep up with report volume produces a portfolio view that is always partially stale — and the positions with the stalest data are, systematically, the positions where current data matters most.

The AI ingestion architecture: three components that each matter

Format-independent extraction capability

A secondaries data ingestion system that requires format-specific mapping for each GP reporting template must be reconfigured every time a GP changes their report format and set up from scratch for every new GP added to the portfolio. This creates maintenance overhead that grows with the portfolio. A system with generalised extraction capability — one that identifies NAV, IRR, MOIC, DPI, TVPI, capital account balance, unfunded commitment, and underlying company list regardless of how they are labelled, positioned, or structured in the source document — scales without proportional maintenance overhead. The technical investment in generalised extraction is significantly greater than in format-specific mapping; the long-term operational cost is significantly lower.

Exception flagging and human review workflow

AI extraction is accurate but not infallible. Extractions that the AI is uncertain about — unusual formatting, ambiguous metric labels, values that fall outside expected ranges for a fund of this type and vintage — should be flagged for human review rather than accepted automatically. The exception rate in a well-calibrated system is low (typically under 5% of extracted data points), but the consequences of passing incorrect data through without review can be significant. Best-in-class systems flag exceptions, route them to a human reviewer through a structured workflow, and track resolution for audit purposes.

Automated cross-validation against prior-period data

Every extracted metric should be compared automatically against the prior-period value for the same fund. Large period-over-period movements — NAV changes that are inconsistent with reported portfolio performance, distributions that do not match expected capital activity, MOIC changes that imply implausible returns for the period — should be flagged for analyst verification before they enter the portfolio data model. This automated cross-validation catches extraction errors that the primary AI might miss and GP reporting anomalies that are important signals for portfolio management.

From data ingestion to portfolio intelligence: what becomes possible

Once GP report data is normalised into a unified, structured, continuously updated dataset, three categories of secondaries portfolio intelligence become achievable that are not feasible through manual processing.

Cross-portfolio underlying company concentration analysis

A secondaries portfolio with 100 underlying fund positions may hold exposure to the same underlying company through five or six different fund positions — positions in funds managed by different GPs, with different entry vintages, representing different portions of the same company’s capital structure. This concentration is invisible when each GP report is processed individually. It becomes visible — and manageable — when all underlying portfolio company data is aggregated into a single queryable dataset. For a secondaries manager managing LP reporting to institutional LPs who want underlying company exposure breakdowns, this capability is a reporting requirement, not just an analytical one.

Performance attribution by vintage, strategy, and GP quality

Disaggregating secondaries portfolio returns by the vintage year of each underlying fund position, the underlying strategy (buyout, growth, credit, infrastructure), and the quality tier of each GP requires consistent, comparable data across all positions. Manual processing — which produces inconsistently structured data from heterogeneous sources — makes this analysis imprecise at best and impossible at institutional portfolio scale. AI-normalised data makes it straightforward.

NAV quality monitoring and GP mark assessment

Systematically tracking GP-reported NAVs over time and comparing them to comparable funds and underlying portfolio company performance requires a database of normalised GP-reported data spanning multiple reporting periods and multiple fund positions. This database exists, in a consistent and queryable form, only if the data ingestion process has been systematic and the data model has been consistently structured. On that foundation, AI can identify mark quality signals — systematic mark stability through volatile periods, persistent deviation from comparable fund marks, mark movements inconsistent with underlying portfolio performance — that are invisible to manual analysis at the scale and frequency needed to be actionable.

A Note on Pepper’s Approach

Pepper ingests GP reports in any format and extracts consistent, structured data automatically, with exception flagging and automated cross-validation. Cross-portfolio concentration analysis, performance attribution, and NAV quality monitoring all operate on the normalised, unified data that ingestion produces. The data ingestion architecture is the foundation on which secondaries portfolio intelligence is built — and it was designed specifically for the scale and format heterogeneity of the secondaries use case.

Related articles

Aren't you just a little curious?

Sign up for our newsletter to receive biweekly updates on the world of asset management, delivered straight to your inbox.