Swartberg Capital is a systematic research and trading firm in South Africa. We build one portfolio from nine strategy sleeves, each grounded in published research and each drawing on a different economic link, and we set the weight of each sleeve by the regime the market is in.
Volatility-targetedCorrelation-budgetedWalk-forward validatedDeflated-Sharpe gatedRegime-conditionedPoint-in-time data
Illustration. A stylised market path, not market history. Each regime has its own drift and volatility: calm growth rises steadily, rising inflation trades sideways with more variation, stress falls quickly as volatility rises, recovery regains part of the fall with volatility still high, and calm growth then resumes below the previous peak. The book above follows the cycle and adjusts its sleeves as conditions change; pause it, or click a regime, to hold one.
These are the styles of strategy we research and run. Each is documented in the academic record, some for close to a century, and each rests on a reason the return should persist: a counterparty who is slower, constrained, buying insurance or pricing the same exposure differently elsewhere. Our work is the construction of each signal, its sizing, and how the sleeves combine.
Every strategy has extended periods of weak returns. Styles that draw on different economic links tend to have them at different times, so the combined book is steadier than any one of its parts. Strategies that move together add leverage, not diversification, which is why we measure the correlation between sleeves every day. In numbers: with N sleeves of equal Sharpe ratio and average pairwise correlation ρ̄, the book's Sharpe ratio scales by √(N / (1 + (N − 1)ρ̄)). Correlation, not the number of strategies, sets the benefit.
Illustration. The regime sequence is an example and the strand widths follow the default switching table further down. It is not a forecast and it shows no performance. Thick: sleeve on. Medium: half size. Faint: off.
A strategy without a published thesis does not enter the book.
The book is managed with a small set of formulas from the published literature. They determine how much evidence a strategy requires, how large it may be, and what diversification contributes.
The annualised Sharpe ratio: mean excess daily return over its volatility. The common unit for every comparison in the book.
Sharpe (1966, 1994)Evidence grows with the square root of time. At a Sharpe ratio of 1, a t-statistic of 2 needs about four years of data; at 0.5, sixteen. New factors should clear t > 3.
Harvey, Liu and Zhu (2016)The probability the true Sharpe ratio is positive after correcting for skew γ3, kurtosis γ4, and SR0, the highest result expected by chance from the number of trials run.
Bailey and López de Prado (2014)For N sleeves of equal Sharpe ratio and average correlation ρ̄. Nine uncorrelated sleeves triple the ratio; at ρ̄ = 0.5 the gain is only a third. Try it below.
Grinold and Kahn (2000); Carver (2015)Each sleeve is sized to a common volatility target σ*, using an exponentially weighted variance estimate that reacts to changing markets.
Moreira and Muir (2017); RiskMetrics (1996)Book volatility from the covariance matrix Σ, shrunk toward a structured target. Each sleeve's risk contribution is measured and held near equal.
Ledoit and Wolf (2004); Maillard, Roncalli and Teïletche (2010)Costs are measured in Sharpe units: commission plus slippage over the annualised volatility of the instrument. No sleeve may spend more than a third of its expected return on trading.
Carver (2015), Systematic TradingThe growth-optimal Kelly fraction overstates size whenever μ is estimated with error, which is always. We cap at half of it.
Kelly (1956); Thorp (2006)Expected shortfall: the average loss on the worst 5% of days, estimated on the bootstrapped distribution. It sees the shape of the tail that value at risk ignores.
Acerbi and Tasche (2002)Multiplier on a single sleeve's Sharpe ratio, assuming equal Sharpe ratios and equal risk. Arithmetic, not a forecast.
A backtest describes history; it does not establish a return. We accept a result only after it has survived tests designed to reject it: anchored walk-forward analysis and purged k-fold cross-validation with an embargo, permutation tests against shuffled data, combinatorially symmetric cross-validation to estimate the probability of backtest overfitting (PBO), and a block bootstrap of the outcome distribution we size for. Most candidates do not pass. That is the purpose of the standard.
Parameters are fitted in sample and tested on the next out-of-sample window, six times over. Observations whose labels overlap the test window are purged, with an embargo after it, so no information leaks across the boundary (López de Prado 2018). The last quarter of history is sealed and touched once.
A permutation test: the same rules on shuffled returns, thousands of times, trace the null distribution luck alone produces. The empirical p-value is the share of shuffled runs that match ours; we require p < 0.05.
A stationary block bootstrap resamples history in blocks, keeping volatility clustering and autocorrelation intact, and gives a distribution of maximum drawdowns. Size and the book's stop come from its tail (the 95th percentile and expected shortfall), not from the one path the backtest happened to take.
The firm's current thresholds, as recorded in our validation file and reviewed for each sleeve before launch. They are standards, not results.
| Test | Hurdle | Why it is there |
|---|---|---|
| Deflated Sharpe ratio (DSR): probability the true Sharpe ratio is above zero, after trials, skew and kurtosis | at least 95% | The strongest of many variants is overstated by construction; deflation removes that effect before the result is assessed. Trials are counted as clusters of similar configurations, and we select the centre of the strongest cluster rather than its single strongest member. |
| Cost speed limit (Carver 2015) | at most ⅓ of pre-cost Sharpe | Turnover times round-trip cost, in Sharpe units. A return that must give up more than a third of itself to trading costs is too fragile to run. |
| Walk-forward windows positive after costs | 4 of 6 | One strong period can lift an average; the result must hold in most periods. |
| The combined book | t of 3 or more | Sleeves that pass are combined at equal risk, and the book must clear a higher bar than any single sleeve (Harvey, Liu and Zhu 2016). |
| Sealed holdout | last 25% | History the research never saw, used once. |
| Permutation test against shuffled returns | p < 0.05 | The result must beat what luck alone produces on the same rules. |
| Profit at twice the dealing costs | still positive | Executed prices are worse than quoted ones; a result that depends on perfect execution is not one we rely on. |
| Parameter stability under ±20% perturbation | 70% of neighbours | Settings 20% either side must also hold, or the result reflects fitting rather than a durable effect. |
| In-sample to out-of-sample decay | at most 50% | A large drop out of sample is the signature of overfitting. |
| Profit factor · tail ratio | 1.3 · 1.0 | Gains must outweigh losses, and the largest losses must not outweigh the largest gains. |
| Drawdown consistency | within the 95th percentile | The worst drawdown must be no deeper than a strategy with the same Sharpe ratio and record length would expect (Magdon-Ismail et al. 2004). A fixed return-to-drawdown threshold penalises long records; this test addresses the question directly. MAR and Sortino are reported on every strategy. |
| Sample size | 200 trades | Fewer trades cannot separate skill from noise; the standard error of a Sharpe estimate shrinks only with the square root of the sample. |
| Probability of backtest overfitting (PBO) | reported | How often the in-sample winner ranks below the median out of sample across combinatorial splits (Bailey, Borwein, López de Prado and Zhu 2017). |
| Paper drift flag | 50 trades, 2 weeks | If paper results fall below the validated range for that long, the candidate goes back down the ladder. |
Our research process runs this ladder continuously, under rules the principals have written down. Artificial intelligence gives the principals the capacity to take every candidate through every test. A candidate advances one rung at a time and earns each test by passing the one before. Most do not progress. Those that reach the top compete for their sleeve's allocation, and the holder is replaced only by a challenger with clearly stronger evidence.
Illustration of the process. Each square is a candidate, coloured by style group; an outlined square has not passed a test and returns to research.
Group challengeMan AHL's description of its research culture, in place of groupthink
The rules are written down before the test. Parameters come from the first part of the history only, and the rest is used once, walking forward.
Selecting the strongest of many variants biases the result upwards. Every result carries the number of configurations tried, a deflated Sharpe ratio and the shuffled-data test.
Evidence builds with the square root of time: separating two good strategies takes years, not a quarter. A challenger takes the allocation only when its evidence is clearly stronger over a long period.
Our systems test and report. No automated process places an order. Going live, and every change of allocation on live capital, is a decision of the principals, recorded with its evidence.
News, fundamentals, market structure and macroeconomic series are brought together each morning. Published rules set the regime on three axes, so any call can be reproduced from the log. Artificial intelligence gives the principals capacity: it reads the day's sources and writes a cited briefing on what changed and why. A two-state hidden Markov model on index returns runs alongside as a cross-check (Hamilton 1989), and every series is stored point-in-time, as first published, so no backtest sees data before it existed. The principals approve the rotation, and the day is declared at 06:00: which sleeves run, at what size and in which direction. Any change during the day is a logged override, scored against the rule. Every scheduled job reports its last good run in the same morning brief, and one that falls silent is flagged the next day.
Regimes persist, then switch abruptly (Hamilton 1989; Ang and Timmermann 2012). We model the state as a hidden Markov chain: each day the market stays in its regime or moves to another with probability pij, and returns are drawn from that regime's own distribution. The probabilities below are illustrative; ours are estimated by maximum likelihood (the EM algorithm) on point-in-time data.
Synthetic example. A market path simulated from the four-state model above, and the Hamilton filter's probability of each regime computed from its daily returns alone, the way the engine reads live data. It is not market history and shows no strategy performance. Hover for the day's probabilities.
Select a regime to see it across the page. Hover a cell for the published reason. The engine computes eight states; four are shown. Our own tests replace this default as they are completed.
Exposure steps down as the book falls from its peak and rebuilds only once the fall has recovered past the step by two points (dashed). The levels are the firm's written plan, not a forecast.
Every active sleeve is scaled to the same ex-ante volatility, wi = σ* / σ̂i, with σ̂ from an exponentially weighted estimator. No sleeve can dominate the book's risk budget.
The covariance matrix is re-estimated daily with Ledoit–Wolf shrinkage, which steadies the estimate when history is short. When two sleeves start to co-move, one is cut back. One month of correlation is a prompt for review, not a verdict.
No sleeve spends more than a third of its pre-cost Sharpe ratio on costs, which caps annual turnover at (SR/3) / c, where c is the round-trip cost in Sharpe units (Carver 2015). It favours daily and weekly horizons over minutes.
If we cannot state who is on the other side of a position and why the return should persist, we do not hold it. Every sleeve's thesis is in the published record.
Mispricings earn the most per unit of risk, hold the least capacity and close the soonest. They sit above the core of the book, never in place of it.
Size and the crash budget come from the 95th percentile of thousands of resampled drawdowns, never the median, and never above half the Kelly fraction. Expected shortfall at 95% is reported daily.
Risk steps down to 75% at a 5% fall from the book's peak, 50% at 10%, and new risk stops at 15% until the principals decide. It rebuilds only after a genuine recovery, so that no decision is improvised during a drawdown (Grossman and Zhou 1993).
Risk is halved the day before and the day of scheduled gap events: central-bank decisions, elections, referendums. A sleeve whose return arises from the event itself is exempt.
Research, paper trading and live trading run on exchange and index data. We do not use prices set by a broker that is also the counterparty to the trade.
That is how the book is managed.