Skip to content
Tim Frenzel

// Insight

Mimicking finance: a 1997 model reproduces 71% of what your fund manager does

13 min read
asset-managementmachine-learningfactor-investingquant-methods

Hand a 1997 sequence model 20 quarters of a mutual fund’s regulatory filings and nothing else. No conversations with the manager, no research notes, no house view on rates. It will call the direction of that fund’s next quarter of position changes 71% of the time. Cohen, Lu and Nguyen ran that experiment across 1,706 actively managed US equity funds and 5,434,702 fund-security-quarter observations from 1990 to 2023. Their framing is a labour-market one, that agents should not be compensated for behaviour a machine can reproduce cheaply, in a market they size at 54 trillion dollars. The finding underneath the headline is narrower and more useful. The residual is where the returns live.

I want to take this one slowly, because the number is the kind that gets quoted into a board deck within a week of publication. It happens to survive scrutiny. What does not survive is the leap the authors make from it.

What the model actually predicts

Three details decide how far 71% travels. The abstract carries none of them.

The target is a three-class label, computed at the fund-security-quarter level. The paper takes the change in share count between consecutive quarters, scales it, and buckets the result. A move of −1% or worse is a sell. A move of +1% or better is a buy. Anything inside that band counts as no material change. So the model is picking one of three outcomes for every position in every fund in every quarter.

The second detail is that these are quarter-over-quarter holdings deltas, reconstructed from N-30D, N-CSR, N-Q and N-PORT filings. The paper uses the word trades throughout. Filings show a position at two points in time, which is a different object. I come back to why that matters below.

The third detail is the model itself. The abstract opens by invoking frontier advancements in artificial intelligence. The architecture is a single-layer LSTM, the recurrent network Hochreiter and Schmidhuber published in 1997. Training runs on 28-quarter rolling windows split chronologically, 20 quarters to fit and the final 8 to test, with no shuffling. That design is clean. The test data genuinely sits in the future relative to the training data, which is more than a lot of machine-learning finance manages.

I read the model choice as the most quotable fact in the paper. The thing reproducing three-quarters of professional asset-management behaviour is a 29-year-old architecture that any competent graduate student can implement in an afternoon.

The lift is real

My first question of any classification result with an uneven label distribution is what a lazy guess would score. If most fund-quarters land in the no-change bucket, a model that always predicts no change would look strong while learning nothing.

The paper answers this directly. The answer is the reason I trust the result. Its Zero Rule benchmark takes each manager’s single most frequent past action and applies it to every position in the test period. That naive rule scores 52%. The LSTM scores 71%, a relative gain of 36.5% at p below 0.0001.

Two ways to guess next quarter's position change
Zero Rule baseline, 52%most frequent past actionapplied to every quarter44.38% below 50%LSTM on filings, 71%20 quarters of filingspredict the next 811.11% below 50%
Both run on the same panel, 1,706 funds and 5,434,702 fund-security-quarter observations from 1990 to 2023. The lift is a 36.5% relative gain, significant at p below 0.0001. The amber panel is the model. Percentages in the third box are the share of managers whose behaviour is predicted at worse than coin-flip accuracy.

The distributional version of that statistic lands harder than the mean. Under the naive rule, 44.38% of managers are predicted at worse than coin-flip accuracy. Under the model, that collapses to 11.11%. The whole distribution tightens and shifts right. The effect reaches well beyond a handful of mechanical index-huggers dragging an average around. Predictability is close to universal across the active universe, peaking at 75% for Mid-cap Blend funds.

Who gets predicted

The cross-section is where the paper gets genuinely interesting. The regressions carry quarter-year and fund fixed effects with errors clustered by quarter.

What makes a manager easy to predict
More predictablelonger tenuremore mandatesolder fundsales limitsLess predictablemanager stakehigh turnoverlarge fundrival funds
Directional signs from fund-level regressions with quarter-year and fund fixed effects, standard errors clustered by quarter, 53,878 fund-quarter observations. Mean manager tenure carries 0.0006 at a t of 9.30, turnover carries -0.0149 at a t of 9.712. The manager-stake result is the one worth arguing about, since a manager with more of her own money in the fund is measurably harder to anticipate.

Seasoned managers are easier to predict. So are managers running more funds and more styles. The authors read that as standardisation under time pressure, which fits. Spread one person across 5 mandates and common processes start doing the work.

The result I keep returning to runs the other way. Manager ownership reduces predictability, with a negative and highly significant coefficient. Managers with more of their own capital in the fund behave in ways the model cannot anticipate from their history. Competition does the same thing, both within a Morningstar category and within a fund family. Predictability, on this evidence, looks like what happens when nobody is pushing.

Fund-level characteristics tell the same story. Larger funds, heavier flows, higher turnover and higher management fees all come with lower predictability, at t-statistics running from 2.2 to 9.7. Funds imposing sales restrictions run the other way, at 0.0578 with a t of 4.874. Note the turnover coefficient of -0.0149 in particular. It matters for the objection I raise below.

The same exercise at the security level tells a consistent story. Smaller positions inside a portfolio are more predictable, with a coefficient of 0.0020, which fits an intuition any PM will recognise. The tail of a book gets managed by rule while the top of it gets managed by attention. Larger firms are harder to predict, so are stocks held by many funds, at a coefficient of −0.0002 with a t of 4.659. Breadth of ownership brings heterogeneous motives, from style rebalancing to benchmark tracking to liquidity needs. That heterogeneity is what defeats the model. Older and more levered firms are easier.

Predictability pays, in four directions

Here is the part a quant should care about. The paper stacks up four independent tests that all point the same way.

Four tests, one direction
Fund quintilesQ1 +0.36%Q5 -0.42%spread 0.79%Factor alphasFF3 +0.37%FF5 +0.37%spread 0.78-0.82%Inside one bookhard 7.98%easy 7.07%gap 91bpsStock long-shortQ1-Q5 424bpsFF5 alpha 386bps
Q1 is the least predictable, Q5 the most. Rows 1 and 2 are 4-quarter cumulative benchmark-adjusted returns and factor alphas at the fund level, all quintile spreads with t-statistics above 2.5. Row 3 compares positions inside a single manager's portfolio, annualised, t of 12.41. Row 4 ranks stocks each quarter by how predictable their holders are, equal weighted, t of 5.74 and 4.96.

Sort funds into quintiles by how predictable their trading is. The least predictable quintile earns 0.36% of benchmark-adjusted abnormal return over 4 quarters with a t-statistic of 2.88. The most predictable quintile loses 0.42%. The spread widens with horizon, from 0.23% at 1 quarter to 0.79% at 4. It holds up under three, four and five-factor adjustment at 0.78% to 0.82%.

Most predictable managers, cumulative abnormal return
0.1-0.5category peer-0.091 quarter-0.222 quarters-0.313 quarters-0.424 quarters
Benchmark-adjusted abnormal return in percent for the most predictable quintile of managers, measured 1 to 4 quarters after sorting. The amber line is the category peer benchmark. The 1-quarter reading carries a t of 1.52 and does not clear significance on its own. By quarter 4 the t reaches 2.36. Over the same horizons the least predictable quintile runs from +0.14% to +0.36%. The spread widens from 0.23% to 0.79%.

Then the paper does something cleverer. It splits each manager’s own book into the positions the model called correctly and the ones it missed. Inside a single portfolio, the holdings the model could not predict returned 7.98% annualised against 7.07% for the ones it could. A 91 basis point gap with a t-statistic of 12.41, from a comparison that holds the manager fixed. Whatever is generating that spread lives inside individual position decisions.

Size the fund-level spread against what the same sample charges. Mean expense ratio across these funds is 1.26%, with a median of 1.19%. The 4-quarter gap between the least and most predictable quintiles is 0.79%. A screen computable from public filings sorts managers by close to two thirds of what they charge for the service.

The most tradeable version comes last. Rank every stock each quarter by the average prediction accuracy of the funds holding it, then go long the least predictable and short the most predictable. That spread returns 424 basis points annualised with a t of 5.74. The five-factor alpha is 386 basis points with a t of 4.96, monotonic across quintiles.

Where I would push back

The empirical work is careful. The interpretation runs ahead of it in three places. The first is a measurement problem worth naming plainly.

A filing cannot see a round trip. Consider two managers. One buys a position in January and sells it in February, ending the quarter flat. The other does nothing for 3 months. Both file identical holdings at both quarter ends, both get labelled no material change, and both feed the model its easiest and most frequent class.

That blind spot cuts against the paper’s own reading. Predictability is meant to proxy for mechanical, routine behaviour. A high-turnover manager churning positions inside the quarter can score as maximally predictable, because the artefact and the routine produce the same label. The paper’s returns results are unaffected by this, since the sort works whatever the label means. The labour-market interpretation is what takes the damage.

The paper’s own data then pushes back on how much this matters, and in fairness it pushes hard. If the artefact dominated, heavy traders would look more predictable, because more of their activity would vanish between filing dates. The measured relationship runs the other way. Turnover carries a coefficient of -0.0149 with a t-statistic of 9.712, meaning funds that trade more are harder to predict. So the blind spot is real as a bound on what a filing can see. Its influence on the headline number looks limited.

Predicting a direction is not replicating a job. This is the leap I would push hardest on. The target variable is the sign of a quarterly position change with a 1% dead band. It says nothing about position sizing, entry timing, the price paid, the risk framework, or the decision to hold through a drawdown. A model that calls the direction of your next quarter’s moves has not thereby earned your seat. The distance between those two claims is most of the argument the paper wants to make about compensation.

The tradeable signal has a plumbing problem. Those 424 basis points come from ranking stocks by the predictability of their holders, which requires the holdings. Funds file with a lag measured in weeks after quarter end. The paper does not price that delay, or turnover, or the cost of shorting the most predictably-held names. Anyone tempted by the long-short should read it beside the falsification checklist this blog applied to backtested predictability in July, where the sobering result was how much apparent signal fails a properly specified null.

One more caveat belongs on the record. This is an NBER working paper. NBER states plainly that its working papers have not been peer-reviewed. It has been through a long seminar circuit including the NBER Big Data and AI conference, which is real vetting of a kind. The status remains preprint.

Active Share, 17 years on

The right ancestor here is Cremers and Petajisto, whose 2009 paper in the Review of Financial Studies introduced Active Share as the fraction of a portfolio that differs from its benchmark index. Their finding was that the measure predicts performance, with the highest Active Share funds beating their benchmarks and the lowest non-index funds trailing.

Both papers ask how active a manager really is, using only holdings. Both find the answer forecasts returns. The question changed shape. Active Share asks how different a manager is from her benchmark. This asks how different she is from her own past. The second question turns out to carry information the first one misses, which the paper establishes by running Fama-MacBeth regressions with Active Share included as a control and watching the predictability effect survive.

That is the strongest defence of the result. A manager can look busy against an index while running a process so regular that 20 quarters of filings give away the next 8.

The bottom line

For an allocator, this is a new screen built from data you already hold. Predictability is computable from public filings. It is stable across the sample, survives the obvious control, and sorts funds monotonically on subsequent abnormal returns. Manager tenure and breadth of mandate now read as mild risk factors. Manager ownership points the other way. The paper never tests ownership against returns directly, so treat that chain as my inference across two of its results.

If I were putting this to work on Monday, I would start with the manager-selection version. It needs holdings you already license, a label definition three lines long, plus a model that predates most of the people reading this. Run it on your existing manager panel, rank by out-of-sample accuracy, then look at what sits in the top quintile before the next allocation review.

For anyone building the systematic version, the honest read is that the 424 basis points are a research finding awaiting a cost model. The within-portfolio test at 91 basis points is the more robust result, because holding the manager fixed removes most of what could otherwise explain the spread.

And for the profession, the number that matters is the one nobody put in the abstract. If a 1997 architecture reproduces 71% of position-change directions from filings alone, then the fee conversation is really about the other 29%. The paper’s own evidence says that residual is where the abnormal returns concentrate, which is a more interesting claim than the replacement story wrapped around it. The work left for the field is measuring the residual directly. No filing-based measure will manage that on its own.

A 1997 sequence model reproduces 71% of a fund manager’s quarterly position changes from public filings, against a 52% naive baseline. The positions it fails to predict outperform the ones it calls correctly by 91 basis points inside the same portfolio. Predictability has become a measurable screen for manager selection. Its limits are the limits of what a quarterly filing can see.

Working on AI that needs to ship?

I help funds, fintechs, and data teams take AI from prototype to production.