Predicting post-merger operating performance with interpretable ML
This thesis tests whether public pre-deal firm and transaction information can rank later operating outcomes without confusing fit for foresight. Five outcome definitions, chronological validation, and interpretable attribution keep the modest predictive signal in view alongside its limits.
Framing
One plus one should exceed two
M&A is sold on the claim that one plus one can exceed two through cost, revenue, asset-use, and financing synergies.
- Above the line: synergy
- Below: no synergy, or worse
Some synergyThey partly fit. The merged firm runs a little above the two on their own.
Duplicated overhead, sites and roles are paid for once instead of twice.
Decided byIntegration execution after close
Value of A + value of B · Promised surplus
Asserted at announcement. None of it is observed in the pre-deal record.
Text equivalent
Both firms are already running, and the left of the figure draws each of them as a line the pre-deal record holds. At the deal the two lines are brought together and four promises are attached: cheaper to run, more sold, more from what they own, and cheaper to borrow. Everything to the right of that point is claimed rather than recorded. Whether the promises add or cost depends on how well the two firms actually fit, which is not in the pre-deal record, so the merged line runs above the level the two firms had already reached when they pull the same way and falls back onto it, or below it, when they pull against each other. The pale lines behind the merged line are every outcome this pairing allows. Proportions are schematic and carry no fitted values.
Answer first
Pre-deal information carries a weak but real rank-ordering signal, concentrated in operational and financial channels rather than classic revenue-synergy stories.
Boundary. Predictive performance is weak and depends on the outcome measure: reported test-set R² ranges from 0.0221 for the Healy CFROA baseline to 0.1167 for the asset-turnover alternative.
46 sec
Project film
The thesis in motion
A short visual overview of the research question, model, and screening application before the full empirical story continues.
Loading player controls… Open the project film
Film summary
The film asks whether pre-deal public information can provide an early signal of post-merger operating performance. It shows the financial and operational inputs, the machine-learning step, the expected direction of change, and a demonstrator for comparing potential synergy signals across deals.
Method
The public-data design
A public-report, target-family design uses 29 pre-deal features and chronological validation without future information.
The design follows one audit trail: define which deals can be measured, keep the outcome family intact, restrict the model to pre-deal evidence, then let later time periods grade it.
13.7%of the universe has every outcome-label ingredient
What the label requires
- Pre-deal cash flow and assets for both firms at year t
- Combined cash flow and assets for the merged entity at t+3
- SIC-year median benchmarks at t and t+3
The evidence covers large, listed, well-documented deals; it does not cover the full M&A universe or private SMEs.
One grade is not a report card.
A student is not judged by one subject. For the same reason, the model is not judged by one operating outcome. Each row is a measurable proxy for realized synergy, not synergy itself.
| Proxy | What it captures | Measure | Role |
|---|---|---|---|
| 01Healy CFROAHealy et al. (1992) | Cash-flow efficiency | CFO / assets | anchor |
| 02Asset turnoverGhosh (2001) | Asset productivity | revenue / assets | headline |
| 03Operating ROAKing et al. (2004) | Profitability | EBIT / assets | diagnostic |
| 04Operating marginKing et al. (2004) | Margin discipline | EBIT / sales | diagnostic |
| 05CAPEX intensityDevos et al. (2009) | Investment behaviour | CAPEX / assets | diagnostic |
No single target is a universal merger-success score. Realized synergy is not directly observable, while supervised machine learning needs a target to predict. These are measurable proxies for synergy, not synergy itself, so the full family is reported together to prevent selective reporting.
Supervised learning needs a number to predict. Getting that number wrong is how merger studies accidentally measure the industry cycle, or the accounting, instead of the deal.
Operating cash flow per unit of assets, so a large firm and a small firm can be compared on the same scale.
Healy-style CFROA is the anchor target, not the only one. Four further operating outcomes are estimated with the identical pipeline.
Text equivalent
The Healy-style outcome is built in five steps: cash flow from operations over total assets; pooled across acquirer and target before the deal; adjusted by the SIC2 industry-year median; differenced between the deal year and three years later; and winsorised at the 1st and 99th percentiles.
Acquirer and target public annual reports at FY−1 only
The model tests whether visible pre-deal information carries signal; it does not claim to know how integration will go.
Hover a node for its formula · select a channel to open its list
Five channels, one economic mechanism each. The grouping is what makes the SHAP attribution in chapter 09 readable: in a 29-column model, correlated single-feature rankings are easy to over-read, so channel aggregation reduces brittleness without claiming stability across repeated fits.
Every feature is engineered from the last annual report filed before the deal was announced. Nothing dated after the announcement enters the model.
Text equivalent
The 29 pre-deal features are grouped into five synergy channels: operational (6), financial (9), cost (7), revenue (5), and macro (2). Grouping is what makes the later SHAP attribution readable, because single-feature stories in a 29-column model are brittle.
Knowledge cutoff · end 2018
Older deals train the model, intermediate deals tune it, and the 2019–2022 cohort grades it once. Every prediction uses only information available before the deal.
Results
How to read the results
Spearman ρ asks whether deals are ordered well; held-out R² asks whether outcome magnitudes are predicted accurately.
R² asks whether the predicted outcome magnitude is close to what happened.
Match every pair · negative R² is worse than using the meanSpearman ρ asks whether the model puts deals in a useful order.
The loop compares true order with model order · rank, not distanceText equivalent
R² measures how closely predicted outcome magnitudes match held-out outcomes; it can be negative when predictions are worse than using the average. The paired bars are illustrative rather than thesis observations. Spearman ρ measures whether higher predicted outcomes are ranked above lower ones, regardless of the exact distance between predictions. Its fixed visible domain is Spearman ρ 0.00 to 0.30.
Source: Published thesis and Public thesis defense presentation.
Text equivalent
Healy CFROA: R² 0.0220, Spearman ρ 0.1610; Asset turnover: R² 0.1167, Spearman ρ 0.2761; Operating ROA: R² -0.0550, Spearman ρ 0.2390; Operating margin: R² -0.0070, Spearman ρ 0.1940; CAPEX intensity: R² -0.0290, Spearman ρ 0.1360.
Source: Published thesis and Public thesis defense presentation.
R² of 0.022 sounds like nothing. Is the model broken?
Not necessarily. Post-acquisition performance is unusually difficult to explain from pre-deal filings, but unlike-for-unlike studies cannot supply a numeric pass mark. The defensible reading comes from the held-out metrics, their uncertainty, and the rank-order evidence together.
Context, not a scoring ruler
- M&A performance literatureKing et al. (2004)
Four heavily studied pre-deal moderators explain near-zero variance across the meta-analytic evidence. This establishes difficulty, not a numeric benchmark band.
- Financial-ML precedentGu, Kelly & Xiu (2020)
Strict out-of-sample validation and nonlinear ensembles are method precedents. Equity-return prediction is a different target and is not used as a numerical comparator.
- Why other finance tasks can look easierAmini et al. (2021)
Capital structure has an equilibrium forcing mechanism; one-time integration outcomes do not. Its much higher R² therefore demonstrates a different problem structure, not a threshold this study should clear.
- This thesis · Healy CFROA baselineHeld-out 2019–20220.0220
R² is positive but its bootstrap interval includes zero. The load-bearing result is modest rank ordering: Spearman ρ = 0.161.
- This thesis · asset turnoverHeld-out 2019–20220.1167
The larger R² is target-specific. Removing operating-level features reduces it to 0.0254, making mean reversion part of the interpretation.
Low R² is contextualised, not celebrated. The model is a weak rank-order screen, never a point predictor.
Text equivalent
The literature explains why post-merger prediction is difficult and why strict out-of-sample machine-learning validation is appropriate; it does not provide numerical benchmark bands comparable across different targets. This thesis therefore reports its held-out values directly: the Healy CFROA baseline has R² 0.022 with a confidence interval that includes zero and Spearman rho 0.161, while asset turnover has R² 0.1167 before an ablation reduces it to 0.0254.
Results
What is learnable
Asset productivity is more learnable than the cash-flow anchor, showing that target definition changes the available signal.
Removing starting operating levels erases most magnitude fit, while much of the rank ordering remains.
Predictable movement is not merger causality.
Most of the signal is consistent with mean-reversion risk. Residual signal remains, so the claim shrinks rather than vanishes.
597 held-out deals · full model compared with the levels-removed specification.
Compare all targetsReturn to the shared fixed-axis view
Interpretability
Inside the model
SHAP shows that the baseline model relies most on operational and financial structure, not a causal story.
Allocation analogy: give each channel the share of the prediction it contributes at the margin. That marginal contribution is a model reading, not a claim about cause.
Single-feature SHAP stories are brittle and easy to over-read: correlated ratios trade rank between runs. Summing absolute attribution into five economic channels is more stable and far harder to cherry-pick after the fact.
Channel shares are relative shares of mean absolute SHAP. Aggregation reduces brittleness; it does not convert attribution into causation.SHAP explains the prediction, never the causal mechanism in the world.
Text equivalent
For the baseline Healy model, operational features account for 34.2% of mean absolute SHAP attribution, financial 30.9%, cost 20.3%, revenue 9.7%, and macro 4.8%. Pre-deal accounting reads structure but cannot observe post-deal execution.
Source: Published thesis and Public thesis defense presentation.
Transfer
The practitioner boundary
The practical output is a compass for inspecting pre-deal signals, not a validated SME prediction model or crystal ball.
Bounded pre-deal signalEvidence, not a verdict
What this thesis adds
- Methodological
A chronological, leak-free, target-family ML design for post-merger operating performance.
- Empirical
Predictability is target-dependent; asset-productivity drift is more learnable than CFROA-style synergy, with the mean-reversion boundary explicit.
- Practical
An interpretable, feasibility-gated advisory framework: inspectable pre-deal signals with explicit limits.
Weak rank-ordering signal survives
Magnitude prediction remains weak and target-dependent
Predictive and bounded, not causal
Interactive research instruments
Continue with the model's logic in your hands.
The thesis establishes a bounded ranking signal. These instruments make that boundary inspectable by letting the reader change the ordering, comparison rule, target, and deal inputs.
DealFit ranking instrument
Compare model order, a transparent reversion heuristic, and seed uncertainty across a live acquisition longlist.
Deal screening instrument
Change deal fundamentals and market context to inspect how four thesis targets alter a bounded screening read.
Interactive prototypes built from the thesis' measurement logic. Schematic outputs are explanatory, not fitted predictions, valuations, or investment recommendations.