Research, in Motion.

Predicting post-merger operating performance with interpretable ML

This thesis tests whether public pre-deal firm and transaction information can rank later operating outcomes without confusing fit for foresight. Five outcome definitions, chronological validation, and interpretable attribution keep the modest predictive signal in view alongside its limits.

Expected merger valueStandalone value + synergy gap
V(A)+V(B)+synergy=V(A+B)
The blue block is the promise: value expected above the two firms operating separately. Whether public pre-deal information can anticipate any later operating outcome is the question this thesis tests.
01Framing

Framing

One plus one should exceed two

M&A is sold on the claim that one plus one can exceed two through cost, revenue, asset-use, and financing synergies.

Where the promise comes fromTwo firms running · one asserted surplus · schematic
  • Above the line: synergy
  • Below: no synergy, or worse
They fitThey clash

Some synergyThey partly fit. The merged firm runs a little above the two on their own.

Duplicated overhead, sites and roles are paid for once instead of twice.

Decided byIntegration execution after close

Value of A + value of B · Promised surplus

Asserted at announcement. None of it is observed in the pre-deal record.

Caption: Every channel is a claim about behaviour after the deal closes. That is precisely the part the pre-deal record cannot show, which is what turns the promise into something that has to be tested rather than assumed.
Text equivalent

Both firms are already running, and the left of the figure draws each of them as a line the pre-deal record holds. At the deal the two lines are brought together and four promises are attached: cheaper to run, more sold, more from what they own, and cheaper to borrow. Everything to the right of that point is claimed rather than recorded. Whether the promises add or cost depends on how well the two firms actually fit, which is not in the pre-deal record, so the merged line runs above the level the two firms had already reached when they pull the same way and falls back onto it, or below it, when they pull against each other. The pale lines behind the merged line are every outcome this pairing allows. Proportions are schematic and carry no fitted values.

Answer first

Pre-deal information carries a weak but real rank-ordering signal, concentrated in operational and financial channels rather than classic revenue-synergy stories.

Boundary. Predictive performance is weak and depends on the outcome measure: reported test-set R² ranges from 0.0221 for the Healy CFROA baseline to 0.1167 for the asset-turnover alternative.

0.0221Healy CFROA test R²
0.1167Asset-turnover test R²
0.2761Asset-turnover Spearman ρ

46 sec

Project film

The thesis in motion

A short visual overview of the research question, model, and screening application before the full empirical story continues.

The film presents the public thesis and its demonstrator. Playback begins only when selected.
Film summary

The film asks whether pre-deal public information can provide an early signal of post-merger operating performance. It shows the financial and operational inputs, the machine-learning step, the expected direction of change, and a demonstrator for comparing potential synergy signals across deals.

04Method

Method

The public-data design

A public-report, target-family design uses 29 pre-deal features and chronological validation without future information.

The design follows one audit trail: define which deals can be measured, keep the outcome family intact, restrict the model to pre-deal evidence, then let later time periods grade it.

01 / Build a measurable universe1995–2022 · 3 / 3 required
30,912raw public deals
4,229measurable at t+3

13.7%of the universe has every outcome-label ingredient

What the label requires

  1. Pre-deal cash flow and assets for both firms at year t
  2. Combined cash flow and assets for the merged entity at t+3
  3. SIC-year median benchmarks at t and t+3
Universe definedOnly complete outcome labels pass
Label completeness gateStrict AND rule · illustrative deals
DealPre-deal firmsMerged entity t+3SIC-year benchmarkDecision
Deal APre-deal firmspresentMerged entity t+3presentSIC-year benchmarkpresent3 / 3 · included
Deal BPre-deal firmspresentMerged entity t+3presentSIC-year benchmarkmissing2 / 3 · excluded
Deal CPre-deal firmspresentMerged entity t+3missingSIC-year benchmarkpresent2 / 3 · excluded
Missing features can remain in the model. A missing label ingredient always excludes the deal because the realized outcome can no longer be constructed.

The evidence covers large, listed, well-documented deals; it does not cover the full M&A universe or private SMEs.

02 / Keep the target family intact5 synergy proxies
One gradeHealy CFROAUseful anchor · incomplete verdict
Report cardFive synergy proxiesOne pipeline · five readings

One grade is not a report card.

A student is not judged by one subject. For the same reason, the model is not judged by one operating outcome. Each row is a measurable proxy for realized synergy, not synergy itself.

ProxyWhat it capturesMeasureRole
01Healy CFROAHealy et al. (1992)Cash-flow efficiencyCFO / assetsanchor
02Asset turnoverGhosh (2001)Asset productivityrevenue / assetsheadline
03Operating ROAKing et al. (2004)ProfitabilityEBIT / assetsdiagnostic
04Operating marginKing et al. (2004)Margin disciplineEBIT / salesdiagnostic
05CAPEX intensityDevos et al. (2009)Investment behaviourCAPEX / assetsdiagnostic

No single target is a universal merger-success score. Realized synergy is not directly observable, while supervised machine learning needs a target to predict. These are measurable proxies for synergy, not synergy itself, so the full family is reported together to prevent selective reporting.

Building the outcome variableHealy, Palepu & Ruback (1992)

Supervised learning needs a number to predict. Getting that number wrong is how merger studies accidentally measure the industry cycle, or the accounting, instead of the deal.

assetsCFOassetsABA + Bfirmindustrytt + 3Δ1st99th

Operating cash flow per unit of assets, so a large firm and a small firm can be compared on the same scale.

Healy-style CFROA is the anchor target, not the only one. Four further operating outcomes are estimated with the identical pipeline.

Caption: each step is one operation on the raw cash-flow number. Select a step to see what that operation removes. Outlier handling follows Ghosh (2001).
Text equivalent

The Healy-style outcome is built in five steps: cash flow from operations over total assets; pooled across acquirer and target before the deal; adjusted by the SIC2 industry-year median; differenced between the deal year and three years later; and winsorised at the 1st and 99th percentiles.

03 / Route only visible evidence

Acquirer and target public annual reports at FY−1 only

29pre-deal features
Costasset utilization
Financialcash-ratio diff
OperationalROA gap
RevenueCAPEX intensity
Macrocredit spread
Frozen modelXGBoostfive target-specific predictions

The model tests whether visible pre-deal information carries signal; it does not claim to know how integration will go.

The full input setAcquirer and target public annual reports at FY−1 only
Asset-turnover gap · rev/assets A − rev/assets BROA gap · EBIT/assets A − EBIT/assets BAcquirer operating margin · EBIT / revenue ATarget cash-flow margin · CFO B / revenue BSame 4-digit SIC · dummySame 2-digit SIC · dummyLeverage gap · debt/assets A − debt/assets BCash-ratio difference · cash/assets A − cash/assets BAcquirer cash / sales · cash A / revenue AQuick ratio · acquirer · (CA − inventory) / CLQuick ratio · target · (CA − inventory) / CLStock payment · dummy · cash < 50%All-cash deal · dummyAltman Z · acquirer · modified Z-scoreAltman Z · target · modified Z-scoreRelative asset size · assets B / assets APPE-intensity difference · PPE/assets A − PPE/assets BInventory-turnover gap · rev/inv A − rev/inv BTarget asset utilisation · revenue B / assets BLog deal value · log(deal value USD)Tender offer · dummyFriendly deal · dummyR&D-intensity difference · R&D/assets A − R&D/assets BCAPEX-intensity difference · capex/assets A − capex/assets BIntangible-intensity difference · intang/assets A − intang/assets BRelative size · sales · revenue B / revenue ACross-border · dummyS&P 500 12-month return · trailing t−13 to t−1Credit spread · Moody's Baa − Aaa, t−169752

Hover a node for its formula · select a channel to open its list

Five channels, one economic mechanism each. The grouping is what makes the SHAP attribution in chapter 09 readable: in a 29-column model, correlated single-feature rankings are easy to over-read, so channel aggregation reduces brittleness without claiming stability across repeated fits.

Every feature is engineered from the last annual report filed before the deal was announced. Nothing dated after the announcement enters the model.

Caption: each channel is given an angular wedge proportional to how many features it carries, so the picture stays truthful about where the input set is concentrated.
Text equivalent

The 29 pre-deal features are grouped into five synergy channels: operational (6), financial (9), cost (7), revenue (5), and macro (2). Grouping is what makes the later SHAP attribution readable, because single-feature stories in a 29-column model are brittle.

04 / Let the future grade the model

Knowledge cutoff · end 2018

01Trainteach
02Validatetune
03Test oncegrade once

Older deals train the model, intermediate deals tune it, and the 2019–2022 cohort grades it once. Every prediction uses only information available before the deal.

05Results

Results

How to read the results

Spearman ρ asks whether deals are ordered well; held-out R² asks whether outcome magnitudes are predicted accurately.

Reading the scorecardsHeld-out 2019–2022 test cohort
R² · outcome magnitude
01
02
03
04
05
Actual outcomeModel prediction

R² asks whether the predicted outcome magnitude is close to what happened.

Match every pair · negative R² is worse than using the mean
Spearman ρ · deal ordering
+11%+7%+4%+1%−2%−5%−9%
Best outcomeWorst outcome

Spearman ρ asks whether the model puts deals in a useful order.

The loop compares true order with model order · rank, not distance
Caption: R² grades the distance between each actual and predicted magnitude. Seven unseen deals repeatedly move from their true order through an unranked scatter into the model's order. Spearman reads how faithfully that second ordering preserves the first.
Text equivalent

R² measures how closely predicted outcome magnitudes match held-out outcomes; it can be negative when predictions are worse than using the average. The paired bars are illustrative rather than thesis observations. Spearman ρ measures whether higher predicted outcomes are ranked above lower ones, regardless of the exact distance between predictions. Its fixed visible domain is Spearman ρ 0.00 to 0.30.

Source: Published thesis and Public thesis defense presentation.

One target family · one comparison frameHeld-out 2019–2022
Healy CFROACash-flow efficiency · anchor
0.0220
0.1610
Asset turnoverAsset productivity · headline
0.1167
0.2761
Operating ROAProfitability
-0.0550
0.2390
Operating marginMargin discipline
-0.0070
0.1940
CAPEX intensityInvestment behaviour
-0.0290
0.1360
The axes and row order stay fixed across the results chapters. Highlighting changes; the evidence does not.
Text equivalent

Healy CFROA: R² 0.0220, Spearman ρ 0.1610; Asset turnover: R² 0.1167, Spearman ρ 0.2761; Operating ROA: R² -0.0550, Spearman ρ 0.2390; Operating margin: R² -0.0070, Spearman ρ 0.1940; CAPEX intensity: R² -0.0290, Spearman ρ 0.1360.

Source: Published thesis and Public thesis defense presentation.

Calibrating the numberOut-of-sample R²

R² of 0.022 sounds like nothing. Is the model broken?

Not necessarily. Post-acquisition performance is unusually difficult to explain from pre-deal filings, but unlike-for-unlike studies cannot supply a numeric pass mark. The defensible reading comes from the held-out metrics, their uncertainty, and the rank-order evidence together.

Context, not a scoring ruler

  1. M&A performance literatureKing et al. (2004)

    Four heavily studied pre-deal moderators explain near-zero variance across the meta-analytic evidence. This establishes difficulty, not a numeric benchmark band.

  2. Financial-ML precedentGu, Kelly & Xiu (2020)

    Strict out-of-sample validation and nonlinear ensembles are method precedents. Equity-return prediction is a different target and is not used as a numerical comparator.

  3. Why other finance tasks can look easierAmini et al. (2021)

    Capital structure has an equilibrium forcing mechanism; one-time integration outcomes do not. Its much higher R² therefore demonstrates a different problem structure, not a threshold this study should clear.

These are interpretive anchors, not numerical benchmark bands.
  1. This thesis · Healy CFROA baselineHeld-out 2019–2022
    0.0220

    R² is positive but its bootstrap interval includes zero. The load-bearing result is modest rank ordering: Spearman ρ = 0.161.

  2. This thesis · asset turnoverHeld-out 2019–2022
    0.1167

    The larger R² is target-specific. Removing operating-level features reduces it to 0.0254, making mean reversion part of the interpretation.

Low R² is contextualised, not celebrated. The model is a weak rank-order screen, never a point predictor.

Caption: literature establishes the prediction problem and its limits; only the two target-specific thesis results share the quantitative axis.
Text equivalent

The literature explains why post-merger prediction is difficult and why strict out-of-sample machine-learning validation is appropriate; it does not provide numerical benchmark bands comparable across different targets. This thesis therefore reports its held-out values directly: the Healy CFROA baseline has R² 0.022 with a confidence interval that includes zero and Spearman rho 0.161, while asset turnover has R² 0.1167 before an ablation reduces it to 0.0254.

06Results

Results

What is learnable

Asset productivity is more learnable than the cash-flow anchor, showing that target definition changes the available signal.

Asset turnover · specification checkHeld-out 2019–2022
Test R² · magnitude
Full model0.1167
Without pre-deal levels0.0254
Spearman ρ · ordering
Full model0.2761
Without pre-deal levels0.2327

Removing starting operating levels erases most magnitude fit, while much of the rank ordering remains.

Mean reversion boundary

Predictable movement is not merger causality.

Conceptual trajectory · not fitted observationsPeak performance settles toward its normal operating level.
AnalogyStaged house

Polished for the viewing and presented at its best.

From a 10/10 viewing, the likely movement is back toward an ordinary lived-in state.

Research designHigh-performing firm

Unusually strong pre-deal asset productivity sets a high starting point.

Later performance can drift back toward its industry norm even when no synergy was forecast.

Settling back is not synergy. Ghosh shows why superior pre-acquisition performance and size require a matched benchmark before post-deal improvement is credited to the acquisition. Ghosh (2001).

Most of the signal is consistent with mean-reversion risk. Residual signal remains, so the claim shrinks rather than vanishes.

597 held-out deals · full model compared with the levels-removed specification.

Compare all targetsReturn to the shared fixed-axis view
Healy CFROACash-flow efficiency · anchor
0.0220
0.1610
Asset turnoverAsset productivity · headline
0.1167
0.2761
Operating ROAProfitability
-0.0550
0.2390
Operating marginMargin discipline
-0.0070
0.1940
CAPEX intensityInvestment behaviour
-0.0290
0.1360
10Interpretation

Interpretability

Inside the model

SHAP shows that the baseline model relies most on operational and financial structure, not a causal story.

From allocation analogy to model readingBaseline Healy model

Allocation analogy: give each channel the share of the prediction it contributes at the margin. That marginal contribution is a model reading, not a claim about cause.

Operational34.2%Financial30.9%Cost20.3%Revenue9.7%Macro4.8%
Why channels, not single features

Single-feature SHAP stories are brittle and easy to over-read: correlated ratios trade rank between runs. Summing absolute attribution into five economic channels is more stable and far harder to cherry-pick after the fact.

Channel shares are relative shares of mean absolute SHAP. Aggregation reduces brittleness; it does not convert attribution into causation.

SHAP explains the prediction, never the causal mechanism in the world.

Caption: the allocation tokens align into the same five channel bars, carrying the analogy into the observed shares of mean absolute SHAP attribution.
Text equivalent

For the baseline Healy model, operational features account for 34.2% of mean absolute SHAP attribution, financial 30.9%, cost 20.3%, revenue 9.7%, and macro 4.8%. Pre-deal accounting reads structure but cannot observe post-deal execution.

Source: Published thesis and Public thesis defense presentation.

17Transfer & evidence

Transfer

The practitioner boundary

The practical output is a compass for inspecting pre-deal signals, not a validated SME prediction model or crystal ball.

Decision support boundaryUse the signal to ask better questions

Bounded pre-deal signalEvidence, not a verdict

The contribution is a disciplined prompt for review. It is neither a causal account nor an automated investment recommendation.
Contribution ledger

What this thesis adds

  • Methodological

    A chronological, leak-free, target-family ML design for post-merger operating performance.

  • Empirical

    Predictability is target-dependent; asset-productivity drift is more learnable than CFROA-style synergy, with the mean-reversion boundary explicit.

  • Practical

    An interpretable, feasibility-gated advisory framework: inspectable pre-deal signals with explicit limits.

Weak rank-ordering signal survives

Magnitude prediction remains weak and target-dependent

Predictive and bounded, not causal

Interactive research instruments

Continue with the model's logic in your hands.

The thesis establishes a bounded ranking signal. These instruments make that boundary inspectable by letting the reader change the ordering, comparison rule, target, and deal inputs.

DealFit ranking instrument

Compare model order, a transparent reversion heuristic, and seed uncertainty across a live acquisition longlist.

Open DealFit ranking instrument

Deal screening instrument

Change deal fundamentals and market context to inspect how four thesis targets alter a bounded screening read.

Open Deal screening instrument

Interactive prototypes built from the thesis' measurement logic. Schematic outputs are explanatory, not fitted predictions, valuations, or investment recommendations.