The research record · start to finish
This desk wasn't built by finding a magic strategy. It was built the slow way — by testing everything and burying almost all of it. Over a sustained research campaign we put roughly ~105 distinct strategies through a survivorship-free, cost-honest, out-of-sample gauntlet. The overwhelming majority died. A small handful of durable, risk-adjusted edges survived — and those became the live book. This page is the honest ledger of that work.
How this is estimated (from the box's own timestamps): the desk was touched on ~87 of its 100 calendar days — nearly every day — producing 1,900+ scripts, tools, and research write-ups. Git records 325 commits, but only across the June–July build phase; the entire research arc after that ran in working sessions and scratch scripts that were never committed, so commits undercount the work.
At ~6 engaged hours per active day — and the heavy days ran from early morning to late night — that's on the order of ~520 hours of hands-on build-test-review time over three-plus months.
Put another way: the finished body of work — 129 survivorship-free studies, 6 paper bots plus the HOUSE real-money bot, and a 34,220-name data warehouse — is what a traditional quant analyst would spend well over 1,000 hours (several months full-time) to produce by hand. The machine compressed months of research into weeks; these figures are honest estimates grounded in file and commit timestamps, not a time clock.
Every idea here was built, backtested on survivorship-free data, charged realistic costs, and split into in-sample / out-of-sample with random-entry controls. These are the ones that did not survive — grouped by family, with a few of the marquee autopsies.
Closing-auction size-surprise. After burying conventional intraday twice more in Sep 2026 — first the price signals, then the order-flow / microstructure ideas an outside review said we'd skipped — one signal came back with a pulse. When a stock's closing auction prints unusually large versus its own norm, the next-day move is positive (~+6–8 bp), and — unlike everything in the graveyard — the effect held its size out-of-sample and is executed inside the auction (market-on-close), so the spread that killed every other intraday signal doesn't eat it.
Then we pushed it to the wall — and it fell. It cleared the first gauntlet (ordinary-day, sign-stable across every sub-period and spec), so we powered it up 56 → 146 names and modelled the cost. It faded with more data (t 2.97 → 2.45), it is weakest in exactly the liquid names you could trade (t≈1.1), and the ~+5–7 bp gross is smaller than the cost of capturing it — the signal is "buy names that just printed a surprise-large closing auction," so you'd be trading into the very imbalance, where impact runs 4–10 bp. A real information whisper, but too small to harvest. Honestly closed.
The precise, defensible claim — not that no intraday edge could ever exist, but that traditional price signals, microstructure / cross-sectional information, and auction imbalance were each tested with progressively richer data and realistic execution costs, and none produced a large, durable, retail-executable edge. The closest signal gave a small gross effect that failed to survive scalable OOS testing and implementation cost. No further intraday research — or paid imbalance data — is justified under this framework. More precise information isn't useful when its economic value is smaller than the friction to act on it.
Connors RSI2 mean-reversion — the cleanest survivor. Slow (multi-day), not clever. Powers the Spare Change bot (The Book, which also used it, was retired 2026-09-27).
The only reliable way to beat the index on both return and drawdown: 70/30 growth/gold. Uncorrelated sleeves lift Sharpe and halve drawdowns.
Combining the fixed engines (equal or risk-sized) beats its own component bots (Book / Futbot / Gringotts individually) on Sharpe and drawdown — robustly, across a broad weight plateau, out-of-sample, and every regime. An un-investable −74% engine becomes a holdable −24% portfolio. Honest caveat: on a matched 2019–2026 window the integrated, levered core book (The Ark) still wins on return and Calmar — so the blend is the calmer alternative, not a replacement. Whatever the core, hold it risk-sized, maintain with drift bands (±10–15% per sleeve, not a calendar — ~1–3 rebalances/yr, negligible cost), and locate sleeves by tax (crypto→Roth, futures→deferred, equity→taxable, gold→sheltered from the 34.6% collectibles rate) for free after-tax return.
Go to cash (now to a safe-haven) when volatility spikes — the sole crash signal that survived every test. The desk's circuit-breaker.
Trend-following across assets is crash-positive — the best diversifier when equities fall. Held static (timing it is overfitting).
3× on the index, 2× on a single asset — 4×/5× die to volatility decay. Leverage smooths or amplifies; it does not create edge.
No retail long-only strategy reliably beats SPY on raw return out-of-sample. The real edge is risk-adjusted — a smoother, shallower ride.
The single fact that shapes the entire desk: almost everything you'd own is the same bet. Measured on 13 years of daily returns — and again inside the two stress regimes, because correlations change exactly when it matters. (Live, refreshed weekly, as the heat-map on the Charts page.)
| Asset | vs S&P · full 13y | 2020 crash | 2022 shock | Verdict |
|---|---|---|---|---|
| Nasdaq / US tech | +0.92 – +0.93 | +0.98 | +0.96 | same bet |
| Semiconductors | +0.80 | +0.95 | +0.89 | same bet |
| Single tech (NVDA/AVGO/MU) | +0.55 – +0.64 | +0.80 – +0.90 | +0.75 – +0.86 | same bet in a crash |
| US REITs | +0.69 | +0.93 | +0.82 | same bet |
| Bitcoin | +0.24 | +0.52 | +0.61 | not a hedge — rises in a panic |
| Bonds (TLT) | −0.15 | −0.38 | +0.07 | hedge that broke post-2021 |
| Gold | +0.09 | +0.24 | +0.17 | steady diversifier ✓ |
| Managed futures (DBMF) | +0.17 | +0.56 | −0.41 | crash hedge (slow) ✓ |
| US dollar | −0.12 | +0.13 | −0.53 | rate-shock hedge ✓ |
QQQ↔VGT = 0.97, VGT↔SMH = 0.89. Diversifying across stocks, sectors and tech names is fake diversification — the whole block converges toward 1.0 in a crash.
Gold (≈0 in every regime), managed-futures (goes negative, −0.41, in a slow crash like 2022), and the dollar (−0.53 in the rate shock) — plus cash. That is the entire real-diversifier list.
Stock–bond correlation flipped from −0.32 (2013–19) to +0.13 (2023–26) — why 60/40 and HFEA failed in 2022. And Bitcoin’s correlation rises to +0.6 in a panic: a return engine, not a hedge.
Gold + managed-futures sleeves, the volatility breaker to cash, and leverage only on the diversified index — every design choice is a direct answer to this one table.
| Bot | What it is | Status |
|---|---|---|
| The Ark | The culminating core book — vol-targeted (65%), trend-gated leverage engine (VGT/SMH) + gold + managed futures + Bitcoin, with a safe-haven crash breaker and a margin guard. (Dip sleeve retired Sep 2026.) | Live paper |
| The Book | The un-levered blueprint — gated core + gold + managed futures, safe-haven breaker. The vol-breaker champion. | Retired 2026-09-27 |
| Spare Change | The original RSI2 dip bot — the first validated edge, shipped as a buy-list. | Live paper |
| Vanguard | Buy-and-hold VGT/SMH + gold with a trend gate on the semis sleeve. | Live paper |
| Balanced Futbot | The desk's IBKR futures bot — MNQ / gold / Bitcoin, vol-targeted & gated. (~0.85 correlated with the Ark.) | Retired 2026-09-27 |
| Gringotts | Bitcoin (IBIT) held above its 30-day line; below it, parks in DBMF (managed futures). The Roth sleeve. | Live paper |
| Oracle | A self-directed allocator that hides in gold/bonds/trend during storms. | Live paper |
| Athena 2 | 60% TQQQ/QQQ + 40% Bitcoin with 3 crash filters (vol breaker, 3-day Bitcoin confirm, 2% Nasdaq band); parks in DBMF. Also the Friday dial inside HOUSE’s 401K. | Live paper (since 09-27) |
| HOUSE v2.9 | The real-money bot for the 3 IBKR accounts: Stocks = Ark engine + gold + DBMF (no margin, no Bitcoin); 401K = engine via 3× ETFs + Bitcoin + the Athena 2 dial; Roth = Gringotts. Daily Bitcoin brake, 35% Bitcoin cap. | Real money — go-live Oct 7 |
| Spider-Man | Pre-registered FinBERT news-sentiment trial — run honestly, proven to have no edge, and buried. | Retired |
Updated 2026-09-30: Atlas, Desk Blend, The Book and Balanced Futbot were retired 2026-09-27; today’s real-money answer is HOUSE v2.9. The ranking below is kept as it was written.
Not a backtest — the four-voice council's opinion (Bill the skeptic · Oracle the structuralist · Claude the builder · Athena the return-seeker) on where real capital should actually go, ranked by return-per-drawdown, robustness, and honest diversification value. The engines are at their ceiling and roughly equal on quality, so this is about fit and conviction, not a winner-take-all — and the council does not fully agree, which is the point.
| # | Book | Why BOCA would fund it here | Loudest dissent |
|---|---|---|---|
| 1 | Atlas (retired 2026-09-27) | The flagship. A core-satellite overlay of bots already running — the best return-per-drawdown (Calmar) on the desk, beat the Ark in 76% of Monte-Carlo worlds at the same tail risk. The real-money core. | — |
| 2 | The Ark | The engine Atlas is built on; still wins on raw return + Calmar over a matched window. The culminating book — but runs ~1.5× margin and braces a −40%. | Bill: too concentrated in one long-tech bet. |
| 3 | Vanguard | Set-and-forget growth + gold, un-levered, robust in every regime. If you touch nothing else, hold this. Gold = the free lunch. | Athena: leaves too much return on the table. |
| 4 | Desk Blend (retired 2026-09-27) | The blend beats its own parts on Sharpe and drawdown across a broad plateau — the calmer alternative to the Ark for the risk-averse. | Oracle: Atlas already does this, better sized. |
| 5 | The Book (retired 2026-09-27) | The un-levered blueprint / vol-breaker champion. The honest, holdable baseline the whole desk is measured against. | — |
| 6 | Spare Change | The first validated edge (RSI2 dip in an uptrend). Real but modest — best as a satellite / buy-list, never a core. | Athena: too small to move the needle. |
| 7 | Gringotts | Genuine out-of-sample edge (BTC trend — now a 30-day line) and undefeated in its game — but speculative. Fund it small, locate it in the Roth. | Bill: one bad crypto winter and it's gone. |
| 8 | Balanced Futbot (retired 2026-09-27) | Managed-futures value is real, but the bot is ~0.56 correlated to the tech core — weak diversification in disguise. Hold the sleeve; the true crash hedge is static DBMF (why Atlas V2 swaps it in). | Oracle: keep for the 401k, don't over-weight. |
| 9 | Athena | The highest ceiling on the desk (levered TQQQ + crypto, step-down-to-QQQ gate) — and the highest risk (−53% drawdown for Athena 1; Athena 2 −45%). Fun money, sized so a bad decade can't end the game. | Bill: ranks it dead last — a leveraged bull-market bet. |
| 10 | Oracle | The self-directed meta-allocator rebuilds the desk's own book — that's prudence, not standalone alpha. Best used as a cross-check, not a primary allocation. | Claude: it's a sanity mirror; keep it as one. |
How it maps to real money (HOUSE v2.9): Stocks (taxable) = the Ark tech engine (VGT/SMH) + gold + DBMF, no margin, no Bitcoin; 401K = the engine through 3× TECL/SOXL at 1/3 + gold + DBMF + Bitcoin, plus the Friday Athena 2 dial (0/20/40% of the household); Roth = Gringotts (Bitcoin above its 30-day line, else DBMF). A daily Bitcoin brake and a 35% household Bitcoin cap sit on top. Tiers, plainly: fund the core first (HOUSE v2.9 / Ark / Vanguard), size the satellites small (Spare Change / Gringotts), and treat Athena 2 as the 401K dial and Oracle as a cross-check.
We ran an adversarial "beat this bot" game on every live bot: build a challenger, prove it wins on the desk's own terms (return first, drawdown second), survivorship-free and costed — and all three reviewers must agree. The pattern was decisive: you can't out-signal the validated designs. No new strategy beat any engine. The only improvements that survived were small, free, or diversification tweaks — and one apparent "win" turned out to be a survivorship-bias trap, caught and reverted.
| Bot | Game result | Shipped |
|---|---|---|
| The Ark | The return-stacked managed-futures challenger was "financed beta" — illusory once funded honestly. The real win was a smarter crash response. | Breaker → 50% momentum-picked safe-haven + 50% T-bills |
| Gringotts | Undefeated across 3 rounds — leverage detonates the drawdown, and nothing out-trends the 44-day gate. | Idle cash → T-bills (SGOV; DBMF since 09-26) |
| The Book | Leverage adds zero Sharpe (again); more gold is genuine diversification — it passed a plateau + beta check (not a gold-bull artifact). | Gold 12% → 20% |
| Spare Change | Apparent hold-longer / more-slot "wins" were survivorship-bias artifacts — caught via the memory and reverted. Already optimal. | No change (reverted) |
| Balanced Futbot | Nothing robustly beats the vol-targeted, trend-gated basket; leverage fails the drawdown test in 100% of Monte-Carlo paths. | No change (at ceiling) |
The through-line across all five: the engines are at their ceiling. The only durable improvements found on any bot were the same two free ideas — park idle cash in T-bills (earning ~4–5% on money that sat at 0%) and a little more diversification (gold, or a momentum-picked crash haven). Not one new signal beat a validated design — which means the bots you're running are already the best versions.
In a Sept-2026 escalation, a four-voice council built its single best consensus challenger to the Ark (a faster crash-brake + risk-parity diversifiers). It edged on return but couldn't cut the drawdown — failing the mandate. Five independent attempts, and the Ark still stands.
Getting the Ark ready for the real IBKR accounts forced three honest questions. Where should each piece live for tax? The dip sleeve realizes almost all its gains short-term every year — so it came out of the taxable account. Does it earn its place anywhere? Tested across the whole household, dropping it everywhere added ~1 point a year at the same Sharpe — so it was retired, and its 20% went to the engine. Can the engine run hotter? A dial sweep found only two robust levers (vol-target and Bitcoin weight); raising the vol-target 50% → 65% added ~3 points a year for ~2 points more drawdown, Sharpe unchanged, peak margin still under 2×.
The refined book — engine + gold + managed futures + Bitcoin — backtests at ~50%/yr, −26% max drawdown (2019–26, now charging margin interest). A new margin guard cuts the book to 1× if the account cushion ever drops below 35%. Real money maps as Taxable = Ark (margin), 401K = the same Ark via 3× ETFs at ⅓ weight (same exposure, no borrowing), Roth = Gringotts parking in DBMF instead of T-bills — running in dry-run first. Together: ~56%/yr, −29% worst drop in the 2019–26 backtest, vs 47%/−28% before — but a 30-year Monte Carlo replaying 1999–2026 (dot-com and 2008 included) puts the household's core at ~14%/yr, so the plan is to expect ~14–19%/yr. Its real edge showed up as survival: the smallest chance of being behind after 10 years of any aggressive option (1.6%) and a typical worst drop of −33% vs QQQ's −66%. Honest caveat: 2019–26 was the best tech/Bitcoin window on record; plan on far less.
If no new signal can beat the engines, the last gain left is how you combine them. Atlas is the answer: the Ark as a dominant core (70%), with Balanced Futbot (20%) and Gringotts (10%) as small, uncorrelated satellites — sized so crypto can lift the return without sinking the ship. It built nothing new; it was a paper overlay of three bots then running. Retired 2026-09-27 — HOUSE replaced it.
In a paired Monte Carlo (5,000 ten-year paths), Atlas beat the Ark on terminal wealth in 76% of worlds — ahead at every percentile, including the unlucky downside — at a catastrophic-drawdown tail that's statistically identical. More return, same risk (as long as crypto doesn't go to zero). The improvement the beat-the-bot games couldn't find in a signal, portfolio construction found in the sizing.
Atlas V2 (Sep 2026): a follow-up test asked whether Atlas is truly diversified — and found Futbot is the least-diversifying satellite (~0.56 correlated to the tech core; essentially more tech-trend in disguise). V2 swaps it for a managed-futures sleeve (DBMF — genuinely crash-positive, up +21.6% in 2022 while tech collapsed) and lifts BTC to 20%: 65 / 15 / 20. Same return, shallower drawdown (−27% vs −37%), and the best return-per-drawdown (Calmar) on the desk. It ran head-to-head with V1 on paper; both were retired 2026-09-27 when HOUSE took over.
| Engine | CAGR | Sharpe | Max drawdown |
|---|---|---|---|
| SPY (buy-and-hold) | 13.6% | 0.81 | −34% |
| The Ark (core alone) | 24.5% | 0.94 | −34% |
| Atlas (V1 · 70/20/10) | 29.1% | 1.11 | −37% |
| Atlas V2 (65/15/20 · DBMF) | 29.9% | 1.30 | −27% |
Survivorship-free data. Every backtest ran on a warehouse that includes the losers and the delisted — the single most common way backtests lie.
Prove it, or bury it. Realistic costs on every trade, in-sample / out-of-sample splits, Monte-Carlo stress across thousands of shuffled histories, and a random-entry control on every "edge" to make sure it wasn't just the market.
Pre-registration. The news-sentiment trial was frozen and forward-tested with its kill-criteria written down in advance — so the "dead" verdict couldn't be rationalized away.
Three voices on every idea. The builder, a structural strategist (Oracle), and a professional skeptic (Bill) whose job is to kill fragile ideas. Nothing shipped without all three agreeing — and they killed a lot.
Outside red-team. The record was also stress-tested by an external AI reviewer. When it argued the intraday graveyard had only killed price signals, we ran the missing order-flow / cross-sectional / auction tests on full tape data — which sharpened the conclusion rather than overturning it, and surfaced the one open auction lead.
The value here isn't a secret strategy — it's the ~105 that we proved don't work, so the money never has to find out the expensive way.
What's left standing is small, honest, and durable: buy dips in uptrends, diversify into what's uncorrelated, cut risk when volatility spikes, and never confuse leverage for edge. That's the whole book.
And the last edge wasn't a signal at all — it was allocation: holding the best engine as a core with small uncorrelated satellites (Atlas) beat the best single engine in 76% of Monte-Carlo worlds, at the same risk (Atlas is now retired; HOUSE carries the idea forward).
The real-money answer is HOUSE v2.9: the same survivors, spread across the right accounts for tax, with a daily graded Bitcoin brake and a 35% Bitcoin cap — nearly the same return as the aggressive version with a far smaller worst drop.
The research bots are paper-traded. HOUSE is in its practice run — shadowing the real accounts before any real trade (go-live starts at 50% on Wed Oct 7).
Compiled from 128 survivorship-free research studies conducted for Spare Change Capital. Research bots are paper-traded and HOUSE is in a practice run; backtest windows and forward tracks are documented per-bot on their own pages. Past results — real or simulated — do not guarantee future returns; every strategy here is expected to endure deep drawdowns. This page is a record of process, not a solicitation.