Trading engine
2026 · open source · paper only
v7 · 26–29 Sep · Scoring every candidate · see v8, the current engine →
I built this to test trading ideas on real US market data without placing live trades. It runs nightly on one Linux box and keeps 25 paper portfolios, each with rules fixed before trading starts. Since 29 September a model scores every nightly candidate next to a fixed rule, and a paired test decides at set dates whether the model adds anything. So far no policy has beaten its control on live data.
Figure 01 · The nightly loop · v7
The model scores, code trades, the ledger decides
Sources on the left, one writer, three decision paths, then sizing and fills, then the evidence on the right. Dashed boxes place no orders. Nothing here moves a strategy to real money on its own.
- Data
- Scores and orders
- Fill
- Evidence
Scoring every candidate
Until v7 the AI made one pick a night from five standouts, and it abstained on 92% of them. At that pace, twenty trades would arrive around February 2027, and twenty trades can only detect an edge of about five percent per trade. That is too little evidence to judge anything.
So the unit of evidence changed from a trade to a scored candidate. Each night the model and a fixed rule both score every candidate in a wider universe, and each score is later labelled with what the stock did. A paired test compares the two on the same names and dates, and I only read it at 60, 90 and 120 scored trading days. Three simulator portfolios trade on the scores with identical mechanics: one ranked by the model, one by the rule, and one by the rule with a model veto.
The model still sizes nothing. Code sizes each position from recent volatility, fixes the stop at entry and places a limit order for the open. Before the open the model can look again and cancel an order, never add or resize one, and every cancelled order keeps the fill it would have had, so the cancel decision is scored too. Headlines and intraday movers can now trigger a decision within minutes, but those triggers only watch and place no orders.
Built, but switched off
Next to it sits a challenger lab: other model policies on the same inputs, a reader for filings and earnings releases, factor-neutral statistics, sequential tests and a score-to-weight optimizer. All of it is built and none of it is on. It waits until the new scoring has run a clean first cycle, and a challenger only takes over a portfolio after a sequential test passes and I approve it.
When a simulated order can fill
I wanted to avoid a backtest using a price it could not have traded. The engine enforces the timing rule in one place: an order signalled from the close of day t fills at the open of day t+1, and attempt_fill raises if you ask for anything else. It raises under python -O too, so optimisation cannot remove the timing check.
The fill price is the open moved against you by a half-spread estimated from the sixty-day median dollar volume, plus five basis points a side. The v7 portfolios enter with a limit order at the open instead, which skips the trade when the stock opens too far above the signal close. An order over one percent of that median volume is rejected outright instead of partially filled, so the engine never has to estimate how much would have filled. A missing bar leaves the order pending for three trading days and then rejects it; the engine never fabricates a bar. Dividends are credited on the ex-date from the same corporate-actions table the screen reads.
same-bar fill
$ python -c 'from sim.fills import attempt_fill; ...'ValueError: look-ahead violation: fill_date 2026-09-17 !> signal_date 2026-09-17 # sim/fills.py — the only guard, verbatimif fill_date <= signal_date: raise ValueError(f"look-ahead violation: fill_date {fill_date} !> signal_date {signal_date}")
Writing the test rules in advance
Each strategy starts as a written test plan: the mechanism, the control it has to beat, one primary statistic, a kill criterion, and the total number of trials. All of that is written down before the first signal. I keep the rule fixed after seeing the result and record failed tests alongside the others.
Ten test plans so far. Seven are closed as rejected or inconclusive: a VIX term-structure timer, turn-of-month, sell-in-May, a drawdown throttle, a vol target, a sector cap and a quarterly ETF rebalance. Each failed the pass mark it set up front. The three calendar timers lost to a static exposure-matched control, which keeps the comparison from simply rewarding a different amount of market exposure. Three are still accruing: a sector-momentum portfolio that needs two hundred shared trading days before its kill rule can fire, a 12-1 cross-sectional momentum portfolio measured against an unscreened control, and a forty-Monday test of SPY's open-to-close drift.
Since 28 September new strategy research runs in a private repo against this engine, under the same rule of writing each test down in advance, and its results stay there. The public log holds 103 tests written down in advance, counted conservatively as 139 trials when a result is corrected for how many ideas were tried.
a report that will not peek
$ cat data/reports/experiments/e1-spy-monday-forward.mdNO RESULT YET — 8 of 40 out-of-sample Mondays.Kill criterion: after 40 Mondays, KILL if mean <= 0 or t < 0.5, net of 20 bp.Current standing: mean -0.29%, t -1.37 — would KILL if applied today,which it is not. 32 Mondays to go.
The missing delisted companies
The price store holds only names listed today. Measured against listed-company counts, that is about eleven percent of the companies that existed in 1996, a quarter of 2003 and forty percent of 2014. No 2008 casualty is in it, so a fold that spans 2008 is one in which those names cannot lose money. The bias is not a constant; it grows the further back a window reaches, and every fold table carries its universe size so a reader can weight it.
The walk-forward therefore never reports absolute return as evidence. A stock-picking portfolio is compared with an equal-weight basket of the same screened names, fold by fold, so the bias sits on both sides of the difference. On that comparison, no screen-driven portfolio beat equal weight on any window of three years or more, and the two portfolios that led the live table in September had drawn down eighteen percent inside two months. That result is in the repo. The engine remains paper-only, and the next research steps are bound by the calendar: the point-in-time tables are not deep enough for a fair stock-selection test until 2029 unless I buy a dataset with the delisted names in it.
Versions
The engine has changed shape several times since July. Figure 1 shows v7. Each earlier version has its own page with the diagram as it stood then.
v1
16–27 Jul 2026
Data and the first paper portfolios →
Yahoo daily bars for about 12,200 listed names (Stooq was blocked on day one), a nightly trend screen, and ten paper portfolios filling at the next open.
v2
28 Jul – 3 Aug
More portfolios and a backtest farm →
Six research-based strategies took the count to 16 paper portfolios, a farm replayed every portfolio over past years, and the rules for splits and dividends were settled.
v3
4–17 Aug
Five portfolios let a model veto entries or tune parameters inside fixed bounds, each paired with an untouched twin, and a news analyst wrote a morning brief. I retired all of it on 18 August.
v4
18 Aug – 17 Sep
Ten-fold walk-forward tests, audits of the fill model and the data sources, and frozen forward monitors that can only continue or kill a strategy. An audit on 2 September found seven simulator bugs.
v5
18–24 Sep
The repo went public. A nightly AI agent trades its own paper portfolio through a locked simulator tool, hourly agents watch without placing orders, a ledger scores every agent decision, and Alpaca is not connected yet while SEC EDGAR capture waits for access.
v6
25 Sep
TradingView quotes and bars reached the intraday agents as research input. They never touched prices, fills or orders.
v7
26–29 Sep
The model now scores every nightly candidate next to a fixed rule, three test portfolios trade those scores, and a paired test decides on fixed check dates. A challenger lab is built but switched off, and new strategy research moved to a private repo.
v8
2 Oct 2026 · current
One shared backtest core
Every strategy is now a short specification over one no-peeking evaluator, with native event and portfolio simulation, one evidence path and much faster full-market runs.
- paper portfolios
- 25 active in the simulator
- AI decision paths
- nightly agent since 21 Sep · candidate scoring since 29 Sep
- test portfolios
- model-ranked · rule-ranked control · rule + model veto
- paired test
- model vs rule on the same names · checked at 60, 90, 120 trading days
- entries
- volatility sizing · limit order at the open · pre-open check may only cancel
- built, switched off
- challenger lab · filing reader · text labs · optimizer
- tests logged
- 103 written down in advance · counted as 139 trials
- strategy modules
- 30 · one file each, rules fixed in advance
- data sources
- Yahoo · Nasdaq · FRED · Cboe · FINRA · CFTC · AAII · NAAIM · SqueezeMetrics
- research-only sources
- TradingView quotes and bars · RSS headlines
- not connected yet
- SEC 8-K (awaiting access) · Alpaca IEX · licensed history
- fill model
- next open · spread tier + 5 bp · ≤ 1 % of 60-day volume
- walk-forward
- 10 folds · train 24 mo · validate 12 mo
- store
- DuckDB · one writer · raw responses kept
- api
- 34 local-only routes · reads plus paper orders
- tests
- 4,078 collected
- python
- ~118k lines outside tests · started 2026-07-16
- status
- paper only · MIT · github.com/ong6/trading-engine
Paper trading only
It holds no credentials and connects to no broker, so it cannot move money. The two services only listen on the machine itself and the repo ships no market data. I built it to test whether the ideas hold up under rules I set in advance. So far, none has passed, and the reports in the repo show why.
// NEXT
Agent skills
One home for coding-agent skills, linked into every repo