Research, in Motion.

Predicting post-merger operating performance with interpretable ML

This thesis tests whether public pre-deal firm and transaction information can rank later operating outcomes without confusing fit for foresight. Five outcome definitions, chronological validation, and interpretable attribution keep the modest predictive signal in view alongside its limits.

Optimised for desktop. Complete on smaller screens.

Answer first

Pre-deal information carries a weak but real rank-ordering signal, concentrated in operational and financial channels rather than classic revenue-synergy stories.

Boundary. Predictive performance is weak and depends on the outcome measure: reported test-set R² ranges from 0.0221 for the Healy CFROA baseline to 0.1167 for the asset-turnover alternative.

0.0221Healy CFROA test R²
0.1167Asset-turnover test R²
0.2761Asset-turnover Spearman ρ
Framing01 / 06
01Framing

Framing

One plus one should exceed two

M&A is sold on the claim that one plus one can exceed two through cost, revenue, asset-use, and financing synergies.

Expected merger valueStandalone value + synergy gap
V(A)+V(B)+synergy=V(A+B)
The blue block is the promise: value expected above the two firms operating separately. Whether public pre-deal information can anticipate any later operating outcome is the question this thesis tests.
02Framing

Framing

The observability problem

The integration behaviours that decide success happen after the deal and are not recorded in pre-deal information.

Public data can describe the firms before close. The integration choices that determine realized value happen later.

Observable before close

  • Public annual reports
  • Firm and transaction characteristics
  • Pre-deal operating and financial structure

Unobservable before close

  • Integration execution
  • Management decisions after close
  • Realized synergy
03Framing

Framing

The bounded question

The thesis asks how much of a post-merger outcome is visible before close, and which part remains out of reach.

Can pre-deal firm and transaction information predict post-merger operating performance?

Predict

Test whether visible pre-deal information carries signal about later outcomes.

Do not infer

Explain why integration succeeds or fails, or identify a causal mechanism.

Prediction is not causal explanation. This is a signal detector, not a crystal ball.

46 sec

Project film

The thesis in motion

A short visual overview of the research question, model, and screening application before the full empirical story continues.

The film presents the public thesis and its demonstrator. Playback begins only when selected.
Film summary

The film asks whether pre-deal public information can provide an early signal of post-merger operating performance. It shows the financial and operational inputs, the machine-learning step, the expected direction of change, and a demonstrator for comparing potential synergy signals across deals.

04Method

Method

The public-data design

A public-report, target-family design uses 29 pre-deal features and chronological validation without future information.

The design follows one audit trail: define which deals can be measured, keep the outcome family intact, restrict the model to pre-deal evidence, then let later time periods grade it.

01 / Construct the sample

1995–2022

30,912raw public deals
4,229measurable at t+3

13.7%of the universe has every outcome-label ingredient

What the label requires

  1. Pre-deal cash flow and assets for both firms at year t
  2. Combined cash flow and assets for the merged entity at t+3
  3. SIC-year median benchmarks at t and t+3

The evidence covers large, listed, well-documented deals; it does not cover the full M&A universe or private SMEs.

02 / Keep the target family intact

5 readings of operating performance

01

Healy CFROA

Cash-flow efficiency

CFO / assets
02

Asset turnover

Asset productivity

revenue / assets
03

Operating ROA

Profitability

EBIT / assets
04

Operating margin

Margin discipline

EBIT / sales
05

CAPEX intensity

Investment behaviour

CAPEX / assets

No single target is a universal merger-success score. Realized synergy is not directly observable in the pre-deal data, but supervised machine learning needs a target to predict. These five later operating outcomes are measurable proxies for synergy, not synergy itself, and the family tests whether learnability changes with the proxy definition.

Is the family cherry-picking? The opposite. All five are estimated with the identical pipeline and reported together, strong and weak alike. Publishing the whole family is precisely what makes selective reporting impossible.

Building the outcome variableHealy, Palepu & Ruback (1992)

Supervised learning needs a number to predict. Getting that number wrong is how merger studies accidentally measure the industry cycle, or the accounting, instead of the deal.

assetsCFOassetsABA + Bfirmindustrytt + 3Δ1st99th

Accounting outliers are clipped rather than dropped, so a handful of restatements cannot drive the fit.

Healy-style CFROA is the anchor target, not the only one. Four further operating outcomes are estimated with the identical pipeline.

Caption: each step is one operation on the raw cash-flow number. Select a step to see what that operation removes. Outlier handling follows Ghosh (2001).
Text equivalent

The Healy-style outcome is built in five steps: cash flow from operations over total assets; pooled across acquirer and target before the deal; adjusted by the SIC2 industry-year median; differenced between the deal year and three years later; and winsorised at the 1st and 99th percentiles.

03 / Route only visible evidence

Acquirer and target public annual reports at FY−1 only

29pre-deal features
Costasset utilization
Financialcash-ratio diff
OperationalROA gap
RevenueCAPEX intensity
Macrocredit spread
Frozen modelXGBoostfive target-specific predictions

The model tests whether visible pre-deal information carries signal; it does not claim to know how integration will go.

The full input setAcquirer and target public annual reports at FY−1 only
Asset-turnover gap · rev/assets A − rev/assets BROA gap · EBIT/assets A − EBIT/assets BAcquirer operating margin · EBIT / revenue ATarget cash-flow margin · CFO B / revenue BSame 4-digit SIC · dummySame 2-digit SIC · dummyLeverage gap · debt/assets A − debt/assets BCash-ratio difference · cash/assets A − cash/assets BAcquirer cash / sales · cash A / revenue AQuick ratio · acquirer · (CA − inventory) / CLQuick ratio · target · (CA − inventory) / CLStock payment · dummy · cash < 50%All-cash deal · dummyAltman Z · acquirer · modified Z-scoreAltman Z · target · modified Z-scoreRelative asset size · assets B / assets APPE-intensity difference · PPE/assets A − PPE/assets BInventory-turnover gap · rev/inv A − rev/inv BTarget asset utilisation · revenue B / assets BLog deal value · log(deal value USD)Tender offer · dummyFriendly deal · dummyR&D-intensity difference · R&D/assets A − R&D/assets BCAPEX-intensity difference · capex/assets A − capex/assets BIntangible-intensity difference · intang/assets A − intang/assets BRelative size · sales · revenue B / revenue ACross-border · dummyS&P 500 12-month return · trailing t−13 to t−1Credit spread · Moody's Baa − Aaa, t−169752

Hover a node for its formula · select a channel to open its list

Five channels, one economic mechanism each. The grouping is what makes the SHAP attribution in chapter 09 readable: in a 29-column model, correlated single-feature rankings are easy to over-read, so channel aggregation reduces brittleness without claiming stability across repeated fits.

Every feature is engineered from the last annual report filed before the deal was announced. Nothing dated after the announcement enters the model.

Caption: each channel is given an angular wedge proportional to how many features it carries, so the picture stays truthful about where the input set is concentrated.
Text equivalent

The 29 pre-deal features are grouped into five synergy channels: operational (6), financial (9), cost (7), revenue (5), and macro (2). Grouping is what makes the later SHAP attribution readable, because single-feature stories in a 29-column model are brittle.

04 / Let the future grade the model

Knowledge cutoff · end 2018

01Trainteach
02Validatetune
03Test oncegrade once

Older deals train the model, intermediate deals tune it, and the 2019–2022 cohort grades it once. Every prediction uses only information available before the deal.

05Results

Results

How to read the results

Spearman ρ asks whether deals are ordered well; held-out R² asks whether outcome magnitudes are predicted accurately.

Reading the scorecardsHeld-out 2019–2022 test cohort
R² · outcome magnitude
01
02
03
04
05
Actual outcomeModel prediction

R² asks whether the predicted outcome magnitude is close to what happened.

Match every pair · negative R² is worse than using the mean
Spearman ρ · deal ordering
+11%+7%+4%+1%−2%−5%−9%
Best outcomeWorst outcome

Spearman ρ asks whether the model puts deals in a useful order.

The loop compares true order with model order · rank, not distance
Caption: R² grades the distance between each actual and predicted magnitude. Seven unseen deals repeatedly move from their true order through an unranked scatter into the model's order. Spearman reads how faithfully that second ordering preserves the first.
Text equivalent

R² measures how closely predicted outcome magnitudes match held-out outcomes; it can be negative when predictions are worse than using the average. The paired bars are illustrative rather than thesis observations. Spearman ρ measures whether higher predicted outcomes are ranked above lower ones, regardless of the exact distance between predictions. Its fixed visible domain is Spearman ρ 0.00 to 0.30.

Source: Published thesis and Public thesis defense presentation.

One target family · one comparison frameHeld-out 2019–2022
Healy CFROACash-flow efficiency · anchor
0.0220
0.1610
Asset turnoverAsset productivity · headline
0.1167
0.2761
Operating ROAProfitability
-0.0550
0.2390
Operating marginMargin discipline
-0.0070
0.1940
CAPEX intensityInvestment behaviour
-0.0290
0.1360
The axes and row order stay fixed across the results chapters. Highlighting changes; the evidence does not.
Text equivalent

Healy CFROA: R² 0.0220, Spearman ρ 0.1610; Asset turnover: R² 0.1167, Spearman ρ 0.2761; Operating ROA: R² -0.0550, Spearman ρ 0.2390; Operating margin: R² -0.0070, Spearman ρ 0.1940; CAPEX intensity: R² -0.0290, Spearman ρ 0.1360.

Source: Published thesis and Public thesis defense presentation.

Calibrating the numberOut-of-sample R²

R² of 0.022 sounds like nothing. Is the model broken?

Not necessarily. Post-acquisition performance is unusually difficult to explain from pre-deal filings, but unlike-for-unlike studies cannot supply a numeric pass mark. The defensible reading comes from the held-out metrics, their uncertainty, and the rank-order evidence together.

Context, not a scoring ruler

  1. M&A performance literatureKing et al. (2004)

    Four heavily studied pre-deal moderators explain near-zero variance across the meta-analytic evidence. This establishes difficulty, not a numeric benchmark band.

  2. Financial-ML precedentGu, Kelly & Xiu (2020)

    Strict out-of-sample validation and nonlinear ensembles are method precedents. Equity-return prediction is a different target and is not used as a numerical comparator.

  3. Why other finance tasks can look easierAmini et al. (2021)

    Capital structure has an equilibrium forcing mechanism; one-time integration outcomes do not. Its much higher R² therefore demonstrates a different problem structure, not a threshold this study should clear.

These are interpretive anchors, not numerical benchmark bands.
  1. This thesis · Healy CFROA baselineHeld-out 2019–2022
    0.0220

    R² is positive but its bootstrap interval includes zero. The load-bearing result is modest rank ordering: Spearman ρ = 0.161.

  2. This thesis · asset turnoverHeld-out 2019–2022
    0.1167

    The larger R² is target-specific. Removing operating-level features reduces it to 0.0254, making mean reversion part of the interpretation.

Low R² is contextualised, not celebrated. The model is a weak rank-order screen, never a point predictor.

Caption: literature establishes the prediction problem and its limits; only the two target-specific thesis results share the quantitative axis.
Text equivalent

The literature explains why post-merger prediction is difficult and why strict out-of-sample machine-learning validation is appropriate; it does not provide numerical benchmark bands comparable across different targets. This thesis therefore reports its held-out values directly: the Healy CFROA baseline has R² 0.022 with a confidence interval that includes zero and Spearman rho 0.161, while asset turnover has R² 0.1167 before an ablation reduces it to 0.0254.

06Results

Results

What is learnable

Asset productivity is more learnable than the cash-flow anchor, showing that target definition changes the available signal.

One target family · one comparison frameHeld-out 2019–2022
Healy CFROACash-flow efficiency · anchor
0.0220
0.1610
Asset turnoverAsset productivity · headline
0.11670.0254
0.27610.2327
Operating ROAProfitability
-0.0550
0.2390
Operating marginMargin discipline
-0.0070
0.1940
CAPEX intensityInvestment behaviour
-0.0290
0.1360
Before the dealUnusually high or low asset productivity

Pre-deal levels describe where a firm starts.

Three years laterSome drift back toward the industry middle

Mean reversion can make later movement partly learnable without observing synergy.

Most of the signal is consistent with mean-reversion risk. Residual signal remains, so the claim shrinks rather than vanishes.

597 held-out deals · hollow markers show the levels-removed specification.

Hollow markers remove pre-deal operating levels.

The axes and row order stay fixed across the results chapters. Highlighting changes; the evidence does not.
Text equivalent

Healy CFROA: R² 0.0220, Spearman ρ 0.1610; Asset turnover: R² 0.1167, Spearman ρ 0.2761; Operating ROA: R² -0.0550, Spearman ρ 0.2390; Operating margin: R² -0.0070, Spearman ρ 0.1940; CAPEX intensity: R² -0.0290, Spearman ρ 0.1360.

Source: Published thesis and Public thesis defense presentation.

07Results

Results

Profitability: order without calibration

Operating ROA and operating margin retain modest rank information while failing to predict the magnitude of profitability changes.

Two profitability proxies · one bounded resultHeld-out 2019–2022

The model can place some stronger deals above weaker ones. It cannot reliably say how much profitability will change.

EBIT / assets

Operating ROA

Profit earned from the combined asset base

Test R²-0.0550magnitude failsSpearman ρ0.2390some order survives
EBIT / sales

Operating margin

Profit retained from each unit of revenue

Test R²-0.0070magnitude failsSpearman ρ0.1940some order survives
One target family · one comparison frameHeld-out 2019–2022
Healy CFROACash-flow efficiency · anchor
0.0220
0.1610
Asset turnoverAsset productivity · headline
0.1167
0.2761
Operating ROAProfitability
-0.0550
0.2390
Operating marginMargin discipline
-0.0070
0.1940
CAPEX intensityInvestment behaviour
-0.0290
0.1360
The axes and row order stay fixed across the results chapters. Highlighting changes; the evidence does not.
Text equivalent

Healy CFROA: R² 0.0220, Spearman ρ 0.1610; Asset turnover: R² 0.1167, Spearman ρ 0.2761; Operating ROA: R² -0.0550, Spearman ρ 0.2390; Operating margin: R² -0.0070, Spearman ρ 0.1940; CAPEX intensity: R² -0.0290, Spearman ρ 0.1360.

Source: Published thesis and Public thesis defense presentation.

Useful for ranking and triage, not for forecasting the size of a profitability gain.

Text equivalent

Operating ROA has held-out R² −0.0550 and Spearman ρ 0.2390. Operating margin has held-out R² −0.0070 and Spearman ρ 0.1940. The model therefore retains modest ranking information for both outcomes while failing to predict their magnitudes better than the held-out mean.

Source: Published thesis and Public thesis defense presentation.

08Results

Diagnostic

CAPEX as a diagnostic

CAPEX intensity tracks investment behaviour but does not provide a standalone synergy interpretation.

One target family · one comparison frameHeld-out 2019–2022
Healy CFROACash-flow efficiency · anchor
0.0220
0.1610
Asset turnoverAsset productivity · headline
0.1167
0.2761
Operating ROAProfitability
-0.0550
0.2390
Operating marginMargin discipline
-0.0070
0.1940
CAPEX intensityInvestment behaviour
-0.0290
0.1360
The axes and row order stay fixed across the results chapters. Highlighting changes; the evidence does not.
Text equivalent

Healy CFROA: R² 0.0220, Spearman ρ 0.1610; Asset turnover: R² 0.1167, Spearman ρ 0.2761; Operating ROA: R² -0.0550, Spearman ρ 0.2390; Operating margin: R² -0.0070, Spearman ρ 0.1940; CAPEX intensity: R² -0.0290, Spearman ρ 0.1360.

Source: Published thesis and Public thesis defense presentation.

The same movement can tell opposite storiesViva block model · looping when visible
CAPEX ↓ · cuttingcould be

Efficiency · disciplineremoving duplicated assets

CAPEX ↑ · raisingcould be

Growth investmentcapacity for the combined firm

CAPEX ↓ · cuttingcould be

Underinvestmentstarving the asset base

CAPEX ↑ · raisingcould be

Overinvestmentempire-building

Caption: direction alone cannot identify whether investment discipline or investment failure produced the change. CAPEX remains a diagnostic channel, not a success label.

CAPEX direction is not a universal success label.

Text equivalent

CAPEX can fall through efficiency or discipline, or through underinvestment. It can rise through growth investment, or through overinvestment. CAPEX intensity, measured as CAPEX divided by assets, has held-out R² −0.029 and Spearman ρ 0.136. It is retained as a diagnostic of investment behaviour.

Source: Published thesis and Public thesis defense presentation.

09Interpretation

Interpretability

What SHAP actually does

Every prediction splits into a baseline plus one exact contribution per feature, computable on trees but only ever about the model.

Start with the intuitionPrediction credit, player by player

If a model predicts Netherlands 2–1 Argentina, SHAP asks how much each player contributed to those predicted goals across possible line-ups.

NED2 : 1ARGpredicted · full time

baseline 0.67+player credits 1.33=2 predicted goals

Player
Model feature
Line-up
Feature coalition
Predicted goals
Predicted post-merger outcome

SHAP explains the prediction, not who caused the real match. In the thesis it shows what the model relied on, never what caused the post-merger outcome.

Illustrative attribution values adapted from the selective thesis defense presentation. Each team's baseline and player credits add to its predicted score.
Text equivalent

The model predicts Netherlands 2 and Argentina 1. For Netherlands, a baseline of 0.67 plus 1.33 of illustrative player credits equals 2 predicted goals. For Argentina, the same baseline plus 0.33 equals 1. Players stand for model features, possible line-ups stand for feature coalitions, and predicted goals stand for the predicted post-merger outcome. This allocates prediction credit and does not establish real-world causality.

The one identity that mattersLocal additivity

f(x) = φ₀ + Σᵢ φᵢ

Every prediction decomposes exactly into a baseline plus one contribution per feature. The parts always sum back to the whole.

6.0φ₀ baseline+0.9Operational+0.7Financial+0.5Cost+0.4Revenue−0.1Macro8.4f(x) prediction

Illustrative decomposition · the arithmetic, not a thesis observation

Caption: the contributions sum back to the prediction exactly. That closure is the whole appeal of SHAP: nothing is left in an unexplained residue.
With 29 features, isn’t exact SHAP intractable?TreeSHAP
Naive Shapley29! conceptual orderings8.8 × 10³⁰ feature orderings

The permutation definition averages a feature’s marginal contribution over every possible ordering. Enumerating those orderings directly is not a practical implementation at 29 features.

TreeSHAPExacttree-path dynamic programming

TreeSHAP uses tree-path dynamic programming across the trained ensemble to return exact φ values in polynomial time. It avoids brute-force permutation or coalition enumeration without pretending the computation is a single pass.

Exact, not approximate. But exact about the model, never about the world.

The line I do not crossAttribution ≠ effect

Feature relevance is a property of the model, not of the world.

  • Janzing, Minorics & Blöbaum (2020)

    Attribution is a causal problem; SHAP without a causal model does not deliver causal effects.

  • Sundararajan & Najmi (2020)

    Shapley variants disagree; conditional-expectation SHAP can violate the axioms it is sold on.

Language rule

saythe model weights X

neverX causes Y

Method anchors: Shapley (1953) · Lundberg & Lee (2017) · Lundberg, Erion & Lee (2018)

Text equivalent

SHAP decomposes a single prediction into additive per-feature contributions that sum to the prediction. The permutation definition considers all 29-factorial conceptual feature orderings; TreeSHAP avoids direct enumeration through exact polynomial-time dynamic programming over the trained tree paths. The exactness is computational: SHAP describes model behaviour, not the economic data-generating process.

10Interpretation

Interpretability

Inside the model

SHAP shows that the baseline model relies most on operational and financial structure, not a causal story.

From allocation analogy to model readingBaseline Healy model

Allocation analogy: give each channel the share of the prediction it contributes at the margin. That marginal contribution is a model reading, not a claim about cause.

Operational34.2%Financial30.9%Cost20.3%Revenue9.7%Macro4.8%
Why channels, not single features

Single-feature SHAP stories are brittle and easy to over-read: correlated ratios trade rank between runs. Summing absolute attribution into five economic channels is more stable and far harder to cherry-pick after the fact.

Channel shares are relative shares of mean absolute SHAP. Aggregation reduces brittleness; it does not convert attribution into causation.

SHAP explains the prediction, never the causal mechanism in the world.

Caption: the allocation tokens align into the same five channel bars, carrying the analogy into the observed shares of mean absolute SHAP attribution.
Text equivalent

For the baseline Healy model, operational features account for 34.2% of mean absolute SHAP attribution, financial 30.9%, cost 20.3%, revenue 9.7%, and macro 4.8%. Pre-deal accounting reads structure but cannot observe post-deal execution.

Source: Published thesis and Public thesis defense presentation.

11Interpretation

Interpretability

How to read the SHAP figures

Spread, colour, slope, and scatter each answer a different question about what the model is leaning on.

Beeswarm · one row per featureFour shapes, four verdicts

Horizontal position is the SHAP push on a single deal. Colour is that deal’s own feature value.

  • Wide spread

    Moves predictions a lot, and variably. High but heterogeneous reliance.

  • Tight at zero

    Barely moves any prediction. The model effectively ignores it.

  • Red one side, blue the other

    Consistent direction: high values push one way, low values the other.

  • Red and blue interleaved

    Non-monotonic or interaction-driven. The push depends on something else.

Caption: schematic clouds on a shared axis, drawn to isolate one reading each. Spread is how hard the model leans on a feature; colour is where that push comes from.
Dependence · one dot per dealClean effect versus interaction

x is the raw feature value, y is that feature’s SHAP push for that deal. Colour is a second feature SHAP picks automatically.

Tight band, clear slope

A clean marginal response. Target cash-flow margin slopes down: high values push the prediction down.

Wide vertical scatter

Interaction-heavy. At one x value the push splits top-to-bottom, and colour explains the split.

Colour splitting top-to-bottom at a single x value means interaction. Colour merely following the line means the two features are correlated, not interacting.

Caption: the same axes, two different diagnoses. A tight band is a marginal response; a vertical split at one x value is a second feature intervening. Reading follows Lundberg et al. (2020).

Spread and slope describe the model’s response surface, never a causal dose–response curve.

Text equivalent

In a SHAP beeswarm each row is a feature and horizontal spread is the size and variability of that feature’s push on the prediction, with colour showing the feature’s own value. In a dependence plot each dot is one deal, the slope shows the direction the model moves, and vertical scatter at a fixed feature value indicates interaction with a second feature.

12Boundaries

Limits

Where the model gives up

The tails pull inward: shrinkage in a low-signal fit, with winsorising dampening how extreme the actuals even look.

Illustrative reconstructionHealy CFROA · reported aggregate pattern

The dots reconstruct the tail-compression mechanism from the reported held-out metrics; they are not the 78 held-out observations.

actual outcome →predicted outcome →

perfect calibrationillustrative compressed slope

Primary cause

Shrinkage toward the conditional mean

The ordering survives: the lowest prediction really is a bad outcome and the highest really is a good one. What compresses is the magnitude. With little signal in a high-variance target, the loss-minimising predictor pulls inward, because confident extreme calls would inflate squared error. That is exactly what ρ = 0.161 alongside R² = 0.022 looks like.

ρ 0.161the order survives0.022the magnitude does not

Secondary factor

Winsorising at the 1st / 99th percentile

The target is capped before fitting, so the plotted “extremes” are already clipped values. Winsorising dampens how extreme the actuals look; it does not create the gap. Without it the underestimation would look larger, not smaller.

The claim is modest rank ordering, never calibrated magnitudes.

Caption: switching between the two views holds the actual outcome fixed on the x-axis and moves only the prediction. The tails travel furthest, which is exactly what shrinkage looks like.
Text equivalent

Extreme outcomes are underestimated for two reasons. The primary cause is shrinkage: in a low-R² fit on a high-variance target, the error-minimising prediction is pulled toward the conditional mean. The secondary factor is that the target is winsorised at the 1st and 99th percentiles, so the plotted extremes are already capped; removing winsorisation would widen the gap rather than close it.

13Boundaries

Robustness

Robustness checks

Distress, culture, and rolling-window checks add nuance while retaining a weak, bounded conclusion.

One question, every checkConclusion under change
  1. Main asset-turnover model

    Does the conclusion survive this change?

    A modest held-out signal makes pre-deal asset-productivity drift inspectable.
    Supports bounded conclusion
  2. Levels-removed ablation

    Does the conclusion survive this change?

    R² 0.1167 → 0.0254; much of the result is consistent with mean-reversion risk.
    Weakens claim
  3. Altman Z comparison

    Does the conclusion survive this change?

    One compact financial-health score: many problems, one needle. It is a diagnostic comparison, not a replacement conclusion.
    Boundary remains
  4. Pompe–Bilderbeek comparison

    Does the conclusion survive this change?

    R² 0.022 → 0.030; the richer distress panel improves diagnostics modestly.
    Boundary remains
  5. Culture and alternative-target checks

    Does the conclusion survive this change?

    Country scores are not firm culture; integration happens between firms, not flags. Alternative targets: Healy CFROA R² 0.0220, Spearman ρ 0.1610; Operating ROA R² -0.0550, Spearman ρ 0.2390; Operating margin R² -0.0070, Spearman ρ 0.1940; CAPEX intensity R² -0.0290, Spearman ρ 0.1360.
    Boundary remains
Caption: each comparison tests the same bounded conclusion. The geometry changes from a full bar to a shortened or open marker as evidence weakens or a boundary remains.
Text equivalent

The main asset-turnover model supports a bounded reading. Removing levels weakens that reading, while Altman Z, Pompe–Bilderbeek, culture, rolling-window, and alternative-target checks retain its boundary rather than establishing a stronger claim. Alternative-target evidence: Healy CFROA R² 0.0220, Spearman ρ 0.1610; Operating ROA R² -0.0550, Spearman ρ 0.2390; Operating margin R² -0.0070, Spearman ρ 0.1940; CAPEX intensity R² -0.0290, Spearman ρ 0.1360. The rolling-window rank correlations are 0.184 for 2013–2015, 0.082 for 2016–2018, and 0.161 for 2019–2022. The check supports partial persistence, not full temporal robustness.

Source: Published thesis and Public thesis defense presentation.

Does the signal hold over time?Healy-style CFROA baseline

Learn from the past, score only on the next unseen years, then roll the window forward. Across 2013–2022 the rank signal stays positive but breathes with the deal cycle.

  1. train 1995–2012
    test 2013–2015ρ 0.184Positive
  2. train 1995–2015
    test 2016–2018ρ 0.082The dip
  3. train 1995–2018
    test 2019–2022ρ 0.161Recovers · COVID inside

Expanding-window forward test · chronological split · no look-ahead

This supports partial persistence, nothing more. Robustness across market regimes is not claimed, the final window contains COVID and the post-COVID deal surge, and this test rolls the CFROA baseline, so asset-turnover stability over time remains future work.

Caption: the train window only ever grows backwards in time and the test window is always the next unseen years. The dip in 2016–2018 is reported, not smoothed away.
Text equivalent

Expanding-window forward tests give Spearman ρ of 0.184 for 2013–2015, 0.082 for 2016–2018, and 0.161 for 2019–2022 on the Healy CFROA baseline. The rank signal stays positive across all three windows but varies with the deal cycle, supporting partial persistence rather than full temporal robustness.

14Boundaries

Economics

Why the model leans on cash

Free-cash-flow theory explains the financial channel, and predicts that the sign reverses outside listed firms.

Why do cash and distress features carry so much of the attribution?Financial channel · 30.9% of attribution
  1. 01
    Cash-rich acquirer

    Free cash flow beyond what the firm’s own project set can absorb.

  2. 02
    Agency cost

    Managers with more money than good ideas face weaker discipline from capital markets.

  3. 03
    Overinvestment

    Acquisitions become the place the surplus goes.

  4. 04
    Value-destroying deals

    Cash reserves are empirically associated with worse acquisition outcomes.

  • Jensen (1986)Free cash flow → agency cost → overinvestment.
  • Harford (1999)Cash reserves → more value-destroying acquisitions.
  • Altman (1968)Distress score as a compact financial-health proxy.
The same variable, two meaningsDo not transfer the sign.

Listed firms: idle cash reads negative, as Jensen-style agency slack.

Family-owned SMEs: the same cash reads positive, as a Harford-style precautionary buffer that keeps integration alive.

Caption: the chain is the theory the financial channel is consistent with. It is a reason the model might lean on cash, not evidence that cash caused the outcome.
Text equivalent

Free-cash-flow theory predicts that cash-rich acquirers over-invest, and the evidence associates cash reserves with value-destroying acquisitions. The model picks up this financial-health structure. In SME settings the sign may reverse, because the same cash balance can be a precautionary buffer rather than agency slack.

15Boundaries

Limits

What the finding does not claim

After the target, attribution, and robustness checks, the surviving claim is predictive and bounded, not causal.

Bounded result

What is learnable

Learnable: some pre-deal asset-productivity drift.

The available signal changes with the target definition, and a small residual remains after levels are removed.

Ruled boundary

What it does not prove

Not proven: stronger synergy forecasting.

It does not observe integration execution, establish a causal mechanism, or turn asset turnover into a universal merger-success score.

16Transfer & evidence

Transfer

Why SME transfer is not copy-paste

The inputs can be computed for private-company records; the labelled outcome cannot, so the chapter stops at feasibility.

The inputs compute · the label does not3,110 revenue-reporting Western European SMEs · ORBIS

Computable inputs

Asset-turnover inputs computable99.4%
3,090 of 3,110 firms · operating revenue and total assets both present
EBIT/assets inputs computable96.2%
2,991 of 3,110 firms above €5m revenue
0Labelled SME post-merger outcomesNo deal IDs paired with t+3 outcomes in ORBIS today, which is the missing road test No training target, no test set.

Two of three answers are encouraging. The third decides the chapter: the inputs compute, and there is still nothing to train against.

Which of the 29 inputs transferWeighted by listed-firm gain

94% of listed-firm gain attribution is computable

0% of it is validated · the gate above is still shut

  1. DDirectly transferable14 features · ~47% of gain

    An ORBIS-equivalent field exists and the logic is unchanged: asset-turnover and ROA gaps, revenue size, SIC overlap, market factors.

  2. MTransferable with adaptation8 features · ~30% of gain

    Valid, but the proxy is redefined: Altman Z′ substituting book equity for market cap, quick ratios, deal value.

  3. RProxy replaced5 features · ~17% of gain

    The concept holds but no ORBIS proxy exists today: inventory turnover, R&D, CAPEX, intangibles, cash ratios.

  4. NNot transferable2 features · ~6% of gain

    No SME analogue, so they are dropped: tender-offer and stock-payment flags, since SME deals are cash-financed.

So can you just run the model on ORBIS?No · three reasons
  • 01
    Acquirer cash flips signDirection can reverse

    The listed-firm model reads cash as negative (Jensen overinvestment). For owner-managed SMEs expect it positive (Harford cash slack). The feature transfers; the sign does not.

  • 02
    A constrained distress bundleAltman Z′ is a proxy

    Private-firm Z′ substitutes book equity for market cap in X₄. Retained earnings are unavailable and total liabilities were not exported, so it is a bundle proxy rather than a reproduced score.

  • 03
    EBITDA is not cash flowCFO is thin below €10m

    Operating cash flow has low ORBIS coverage for small firms, so EBITDA margin substitutes, but EBITDA ignores working-capital movement. That makes it a CFROA proxy, not the Healy target.

Caption: the coverage numbers are the reason this chapter exists, and the caveats are the reason it stops at feasibility.

Field-level computability under workable ORBIS conditions is not uniform availability, and construct-level portability is not validated predictive portability. Without a labelled SME outcome there is still no training target and no test set.

Text equivalent

For 3,110 revenue-reporting Western European SMEs, asset-turnover inputs are computable for 3,090 firms (99.4%) and EBIT/assets inputs for 2,991 firms (96.2%), while zero labelled post-merger SME outcomes are available. Grading all 29 features for ORBIS portability leaves roughly 94% of listed-firm gain attribution in the directly-transferable, adaptable, or proxy-replaced classes, but computability is not validation.

17Transfer & evidence

Transfer

The practitioner boundary

The practical output is a compass for inspecting pre-deal signals, not a validated SME prediction model or crystal ball.

Decision support boundaryUse the signal to ask better questions

Bounded pre-deal signalEvidence, not a verdict

Use

Use for ranking and triage

Compare the visible pre-deal profile across a pipeline, then identify cases that merit deeper diligence.

Use

Use for sensitivity analysis

Test whether a provisional view changes when observable assumptions, outcomes, or scenarios move.

Do not use

Do not use as causal proof

A model reliance pattern does not establish why an operating outcome changed after a deal.

Do not use

Do not use as an automatic deal decision

The score cannot see integration quality, management judgement, or the commercial case that remains to be tested.

The contribution is a disciplined prompt for review. It is neither a causal account nor an automated investment recommendation.
Contribution ledger

What this thesis adds

  • Methodological

    A chronological, leak-free, target-family ML design for post-merger operating performance.

  • Empirical

    Predictability is target-dependent; asset-productivity drift is more learnable than CFROA-style synergy, with the mean-reversion boundary explicit.

  • Practical

    An interpretable, feasibility-gated advisory framework: inspectable pre-deal signals with explicit limits.

18Transfer & evidence

Evidence

What the claims stand on

Ten load-bearing claims, each tied to the papers that anchor it and the chapter that defends it.

Claim × anchor work10 claims · 19 anchor works
Claims carried
2
2
2
1
1
1
1
1
1
1
1
1
1
1
1
1
1
1
1

The most-used anchors carry 2 of the 10 claims: Ghosh (2001), Healy, Palepu & Ruback (1992), King et al. (2004). 3 of 19 works support more than one claim; the rest carry one claim apiece.

Caption: one filled square is one dependency. Select a source column to trace its claims in cobalt; outlined cells remain visible for context. The footer bars show dependency concentration without turning citation frequency into evidence quality.

These references support bounded claims. None of them makes the model causal, and none of them turns a feasibility map into a validated SME result.

Text equivalent

Ten load-bearing claims in the thesis are each tied to their anchor papers: target construction (Healy, Palepu & Ruback 1992; Ghosh 2001), target family (Healy, King, Devos), low predictability (King 2004; Gu, Kelly & Xiu 2020; Amini 2021), mean reversion (Ghosh 2001), SHAP method (Shapley 1953; Lundberg 2017, 2018), SHAP limits (Janzing 2020; Sundararajan & Najmi 2020), distress (Altman 1968; Pompe & Bilderbeek 2005), culture (Hofstede), the cash channel (Jensen 1986; Harford 1999), and SME feasibility (Arvanitis & Stucki; Golubov & Xiong).

19Transfer & evidence

Evidence

Evidence archive

The public presentation, repository, and published thesis preserve the evidence and its limits for inspection.

These are the inspectable artefacts behind the results, limitations, and practitioner boundary above. Captions link every visual back to its public evidence.

Dot plot comparing published out-of-sample R-squared and Spearman rank correlation across six post-merger outcome models

Dot plot comparing published out-of-sample R-squared and Spearman rank correlation across six post-merger outcome models. Public thesis extension summary.

Text equivalent

Test R² / Spearman ρ: Healy CFROA 0.0221 / 0.1609; Pompe 0.0296 / 0.1975; culture 0.0120 / 0.1318; Pompe plus culture 0.0304 / 0.2035; asset turnover 0.1167 / 0.2761; asset turnover without level features 0.0254 / 0.2327. Asset turnover is more predictable, but mean-reversion concerns bound the conclusion.

Chronological model validation timeline with training from 1995 to 2015, validation from 2016 to 2018, and held-out testing from 2019 to 2022

Chronological model validation timeline with training from 1995 to 2015, validation from 2016 to 2018, and held-out testing from 2019 to 2022. Public thesis defense deck.

Text equivalent

Chronological split: train on 1995–2015 deals, validate on 2016–2018 deals, and test once on the held-out 2019–2022 cohort. The knowledge cutoff is the end of 2018, so the model never trains on future deals.

Horizontal bars showing operational and financial features account for most mean absolute SHAP attribution in the baseline post-merger model

Horizontal bars showing operational and financial features account for most mean absolute SHAP attribution in the baseline post-merger model. Public thesis defense deck.

Text equivalent

Baseline Healy-model shares of mean absolute SHAP attribution: operational 34.2%, financial 30.9%, cost 20.3%, revenue 9.7%, and macro 4.8%. SHAP describes what the model relied on, not a causal mechanism.

Public record

Sources and repository