Track record

People searching for AI stock picker accuracy usually find a hit rate with no losing weeks and no way to check it. This page is the opposite: every graded week TensorSwing has published under its public protocol, every model including the ones that lost, and a recount recipe at the bottom.

01 The numbers so far

Weeks gradedCalls gradedBest modelMajority vote
2 W36 · W37 1,640 7 scratch · 19 abstain 55.9% PATCHTST · 128 of 229 56.1% 133 of 237 names

The verified chain begins at 2026-W36. W34 was a reduced-schema bootstrap week, and W35 was published to the app without proof coverage during the move from three models to seven; it is excluded from the chain and written up as incident 001 in the public repository. Neither is counted above. A coin-flipping model posts 55.9% over 229 calls about one time in 25, so read the table as a baseline being established, not as a result.

02 Graded calls by model

GRADED CALLS · 2026-W36 – W37 1,640 CALLS · 7 SCRATCH · 19 ABSTAIN
MODELW36W37HIT RATE÷ NAIVE
PATCHTSTTENSORSWING · HOUSE 57.0% 54.8% 55.9% 0.96
GRANITE TTM r2IBM · PRETRAINED 58.1% 50.0% 54.0% 0.97
CHRONOSAMAZON · PRETRAINED 57.6% 45.4% 51.5% 1.01
N-HITSTENSORSWING · HOUSE 61.9% 40.9% 51.3% 1.06
TOTO 2.0DATADOG · PRETRAINED 48.3% 53.8% 51.1% 0.93
TIMESFM 2.5GOOGLE · PRETRAINED 46.6% 49.6% 48.1% 1.05
CHRONOS-2AMAZON · PRETRAINED 46.6% 47.1% 46.8% 0.93
MAJORITY4-OF-7 VOTE · TIES SIT OUT 62.7% 49.6% 56.1% —
HIT RATE = WINS ÷ (WINS + LOSSES), DIRECTION ONLY, BEFORE COSTS. ÷ NAIVE = MEDIAN ABSOLUTE ERROR OVER THE ERROR OF FORECASTING NO MOVE; BELOW 1.00 BEATS GUESSING ZERO. RECOUNTED FROM labels/2026-W36.json AND labels/2026-W37.json.

Hit rate counts direction only. Scaled error (÷ naive) divides each model's median absolute error by the error of forecasting no change at all; below 1.00 beats guessing zero. Every model is within a few percent of 1.00, which is the honest reading of two weeks: no model has yet shown it can size a move better than a flat guess. A scratch is a week whose move rounds to exactly 0.00% and counts as neither; an abstain is a name a model declined to call, sealed as such and opened in the audit.

03 What this page no longer claims

Earlier copy on this page showed a three-model record from W01–W03 with a majority vote measured against SPY. Those weeks predate the public proof chain and cannot be recounted from the repository, so they have been removed rather than kept as unverifiable decoration. A SPY comparison returns here once it can be recomputed from published labels alone.

04 Why a bad week cannot be removed

Every Sunday all 833 calls are hashed with salts into a Merkle tree, and the root is timestamped on Bitcoin via OpenTimestamps and by two RFC 3161 authorities before Monday's open. Friday's labels are graded against that fixed set and stamped again before the audit beacon fires. Dropping a losing call, or a losing week, would break a root that was already public. A missed obligation is written into a hash-chained incident log; the log is part of the record, and the next week's manifest pins it.

05 Check any week yourself

The record, the protocol it is built to, and the verification tooling are public. No account and no app are needed:

$ git clone https://github.com/YermekIbrayev/tensorswing-record.git
$ cd tensorswing-record
$ git checkout 2026-W37
$ python3 scripts/verify.py 2026-W37
$ ots verify 2026-W37.manifest.ots -f manifests/2026-W37.manifest.json

Every week is tagged in the repository; checking out the week's tag pins the verifier to the exact revision published with that week. The table above is a plain recount of labels/2026-W36.json and labels/2026-W37.json: count each model's win and loss entries and you get the same numbers. The procedure itself is normative in §13 of PROTOCOL.md.

What a passing verification proves, what it does not, and what the verifier reports today are on the verify-it-yourself page; the grading rule and the full weekly process are on the methodology page; which 119 instruments the models call is the universe rule.

06 What this record is not

Results shown are historical fact and are not a projection of future results. TensorSwing publishes informational research commentary produced by statistical models without human review of individual calls. It is not investment advice, it is not personalised to you, and no trades are placed in the app. Full risk language: Risk Disclosure.