Skip to content
← Selected projects

Trading engine

2026 · open source · paper only

v7 · 26–29 Sep · Scoring every candidate · see v8, the current engine →

I built this to test trading ideas on real US market data without placing live trades. It runs nightly on one Linux box and keeps 25 paper portfolios, each with rules fixed before trading starts. Since 29 September a model scores every nightly candidate next to a fixed rule, and a paired test decides at set dates whether the model adds anything. So far no policy has beaten its control on live data.

Figure 01 · The nightly loop · v7

The model scores, code trades, the ledger decides

Sources on the left, one writer, three decision paths, then sizing and fills, then the evidence on the right. Dashed boxes place no orders. Nothing here moves a strategy to real money on its own.

  • Data
  • Scores and orders
  • Fill
  • Evidence
Yahoo · Nasdaqdaily + intradayTradingViewresearch onlyRSS headlines11 feedsSEC 8-Kwaits on accessMacro feedsFRED · Cboe1ONE WRITER2pricesverifiedfactsas-of timeheadlinestime-stampedpaper ledgerfills · cashrule baselinefixed rankingmodel scoresevery candidate3nightly agentone locked trade toolevent triggersnews · movers · watch onlychallenger labbuilt · switched offsizing + riskcode only4pre-open checkmay only cancelfillnext open or limit5FILLS · DIVIDENDSLATERIBKR brokernot open yetEVERY DECISION, KEPTledgerlabels at 1–20 days6paired testmodel vs rulefixed check dates60 · 90 · 120 daysKEEP · DROP3 test portfoliosmodel · rule · vetoevery test loggedwins and lossesI DECIDE GO-LIVESOURCESDUCKDBDECIDESIZE · FILLEVIDENCE
fig. 1 — one night in v7. ① Yahoo, Nasdaq, TradingView, RSS headlines and a set of macro publishers feed the collectors; SEC 8-K capture waits on access. ② One writer commits every batch to DuckDB, and each fact carries the time it became available. ③ A fixed rule and a model both score every candidate, and the nightly agent keeps its own portfolio; event triggers and the challenger lab place no orders. ④ Code sizes each position and checks risk; before the open the model may cancel an order but never add one. ⑤ Orders fill at the next open or a limit order at the open, and nowhere else. ⑥ Every decision lands in one ledger, is labelled later, and a paired test against the rule decides at 60, 90 and 120 trading days. It is all still paper, so the IBKR broker link is not open yet.

Scoring every candidate

Until v7 the AI made one pick a night from five standouts, and it abstained on 92% of them. At that pace, twenty trades would arrive around February 2027, and twenty trades can only detect an edge of about five percent per trade. That is too little evidence to judge anything.

So the unit of evidence changed from a trade to a scored candidate. Each night the model and a fixed rule both score every candidate in a wider universe, and each score is later labelled with what the stock did. A paired test compares the two on the same names and dates, and I only read it at 60, 90 and 120 scored trading days. Three simulator portfolios trade on the scores with identical mechanics: one ranked by the model, one by the rule, and one by the rule with a model veto.

The model still sizes nothing. Code sizes each position from recent volatility, fixes the stop at entry and places a limit order for the open. Before the open the model can look again and cancel an order, never add or resize one, and every cancelled order keeps the fill it would have had, so the cancel decision is scored too. Headlines and intraday movers can now trigger a decision within minutes, but those triggers only watch and place no orders.

Built, but switched off

Next to it sits a challenger lab: other model policies on the same inputs, a reader for filings and earnings releases, factor-neutral statistics, sequential tests and a score-to-weight optimizer. All of it is built and none of it is on. It waits until the new scoring has run a clean first cycle, and a challenger only takes over a portfolio after a sequential test passes and I approve it.

When a simulated order can fill

I wanted to avoid a backtest using a price it could not have traded. The engine enforces the timing rule in one place: an order signalled from the close of day t fills at the open of day t+1, and attempt_fill raises if you ask for anything else. It raises under python -O too, so optimisation cannot remove the timing check.

The fill price is the open moved against you by a half-spread estimated from the sixty-day median dollar volume, plus five basis points a side. The v7 portfolios enter with a limit order at the open instead, which skips the trade when the stock opens too far above the signal close. An order over one percent of that median volume is rejected outright instead of partially filled, so the engine never has to estimate how much would have filled. A missing bar leaves the order pending for three trading days and then rejects it; the engine never fabricates a bar. Dividends are credited on the ex-date from the same corporate-actions table the screen reads.

same-bar fill

$ python -c 'from sim.fills import attempt_fill; ...'ValueError: look-ahead violation: fill_date 2026-09-17 !> signal_date 2026-09-17 # sim/fills.py — the only guard, verbatimif fill_date <= signal_date:    raise ValueError(f"look-ahead violation: fill_date {fill_date} !> signal_date {signal_date}")
fig. 2 — the guard as it fails, and the two lines that make it fail. There is no second place a fill can be created.

Writing the test rules in advance

Each strategy starts as a written test plan: the mechanism, the control it has to beat, one primary statistic, a kill criterion, and the total number of trials. All of that is written down before the first signal. I keep the rule fixed after seeing the result and record failed tests alongside the others.

Ten test plans so far. Seven are closed as rejected or inconclusive: a VIX term-structure timer, turn-of-month, sell-in-May, a drawdown throttle, a vol target, a sector cap and a quarterly ETF rebalance. Each failed the pass mark it set up front. The three calendar timers lost to a static exposure-matched control, which keeps the comparison from simply rewarding a different amount of market exposure. Three are still accruing: a sector-momentum portfolio that needs two hundred shared trading days before its kill rule can fire, a 12-1 cross-sectional momentum portfolio measured against an unscreened control, and a forty-Monday test of SPY's open-to-close drift.

Since 28 September new strategy research runs in a private repo against this engine, under the same rule of writing each test down in advance, and its results stay there. The public log holds 103 tests written down in advance, counted conservatively as 139 trials when a result is corrected for how many ideas were tried.

a report that will not peek

$ cat data/reports/experiments/e1-spy-monday-forward.mdNO RESULT YET — 8 of 40 out-of-sample Mondays.Kill criterion: after 40 Mondays, KILL if mean <= 0 or t < 0.5, net of 20 bp.Current standing: mean -0.29%, t -1.37 — would KILL if applied today,which it is not. 32 Mondays to go.
fig. 3 — the Monday experiment refuses to report at eight observations. The report prints what the verdict would be and then says why it is not one.

The missing delisted companies

The price store holds only names listed today. Measured against listed-company counts, that is about eleven percent of the companies that existed in 1996, a quarter of 2003 and forty percent of 2014. No 2008 casualty is in it, so a fold that spans 2008 is one in which those names cannot lose money. The bias is not a constant; it grows the further back a window reaches, and every fold table carries its universe size so a reader can weight it.

The walk-forward therefore never reports absolute return as evidence. A stock-picking portfolio is compared with an equal-weight basket of the same screened names, fold by fold, so the bias sits on both sides of the difference. On that comparison, no screen-driven portfolio beat equal weight on any window of three years or more, and the two portfolios that led the live table in September had drawn down eighteen percent inside two months. That result is in the repo. The engine remains paper-only, and the next research steps are bound by the calendar: the point-in-time tables are not deep enough for a fair stock-selection test until 2029 unless I buy a dataset with the delisted names in it.

Versions

The engine has changed shape several times since July. Figure 1 shows v7. Each earlier version has its own page with the diagram as it stood then.

  1. v1

    16–27 Jul 2026

    Data and the first paper portfolios →

    Yahoo daily bars for about 12,200 listed names (Stooq was blocked on day one), a nightly trend screen, and ten paper portfolios filling at the next open.

  2. v2

    28 Jul – 3 Aug

    More portfolios and a backtest farm →

    Six research-based strategies took the count to 16 paper portfolios, a farm replayed every portfolio over past years, and the rules for splits and dividends were settled.

  3. v3

    4–17 Aug

    AI adjusting the portfolios →

    Five portfolios let a model veto entries or tune parameters inside fixed bounds, each paired with an untouched twin, and a news analyst wrote a morning brief. I retired all of it on 18 August.

  4. v4

    18 Aug – 17 Sep

    Evidence first →

    Ten-fold walk-forward tests, audits of the fill model and the data sources, and frozen forward monitors that can only continue or kill a strategy. An audit on 2 September found seven simulator bugs.

  5. v5

    18–24 Sep

    An agent in the loop →

    The repo went public. A nightly AI agent trades its own paper portfolio through a locked simulator tool, hourly agents watch without placing orders, a ledger scores every agent decision, and Alpaca is not connected yet while SEC EDGAR capture waits for access.

  6. v6

    25 Sep

    TradingView for research →

    TradingView quotes and bars reached the intraday agents as research input. They never touched prices, fills or orders.

  7. v7

    26–29 Sep

    Scoring every candidate →

    The model now scores every nightly candidate next to a fixed rule, three test portfolios trade those scores, and a paired test decides on fixed check dates. A challenger lab is built but switched off, and new strategy research moved to a private repo.

  8. v8

    2 Oct 2026 · current

    One shared backtest core

    Every strategy is now a short specification over one no-peeking evaluator, with native event and portfolio simulation, one evidence path and much faster full-market runs.

paper portfolios
25 active in the simulator
AI decision paths
nightly agent since 21 Sep · candidate scoring since 29 Sep
test portfolios
model-ranked · rule-ranked control · rule + model veto
paired test
model vs rule on the same names · checked at 60, 90, 120 trading days
entries
volatility sizing · limit order at the open · pre-open check may only cancel
built, switched off
challenger lab · filing reader · text labs · optimizer
tests logged
103 written down in advance · counted as 139 trials
strategy modules
30 · one file each, rules fixed in advance
data sources
Yahoo · Nasdaq · FRED · Cboe · FINRA · CFTC · AAII · NAAIM · SqueezeMetrics
research-only sources
TradingView quotes and bars · RSS headlines
not connected yet
SEC 8-K (awaiting access) · Alpaca IEX · licensed history
fill model
next open · spread tier + 5 bp · ≤ 1 % of 60-day volume
walk-forward
10 folds · train 24 mo · validate 12 mo
store
DuckDB · one writer · raw responses kept
api
34 local-only routes · reads plus paper orders
tests
4,078 collected
python
~118k lines outside tests · started 2026-07-16
status
paper only · MIT · github.com/ong6/trading-engine

Paper trading only

It holds no credentials and connects to no broker, so it cannot move money. The two services only listen on the machine itself and the repo ships no market data. I built it to test whether the ideas hold up under rules I set in advance. So far, none has passed, and the reports in the repo show why.

Ong Jun Xiong

SOFTWARE ENGINEER · SINGAPORE

ContactHobbiesArchiveNotesUI PackGitHubLinkedInSource

© 2026 Ong Jun Xiong