Track record
Two graded weeks and 1,640 calls is not evidence of anything. It is shown anyway, in full, including the week five of seven models finished under 50%. Direction and size are scored separately, every call is graded against Friday's close before costs, and every number here traces to a label file whose hash was public before the week began.
People searching for AI stock picker accuracy usually find a hit rate with no losing weeks and no way to check it. This page is the opposite: every graded week TensorSwing has published under its public protocol, every model including the ones that lost, and a recount recipe at the bottom.
01 The numbers so far
| Weeks graded | Calls graded | Best model | Majority vote |
|---|---|---|---|
| 2 W36 · W37 | 1,640 7 scratch · 19 abstain | 55.9% PATCHTST · 128 of 229 | 56.1% 133 of 237 names |
The verified chain begins at 2026-W36. W34 was a reduced-schema bootstrap week, and W35 was published to the app without proof coverage during the move from three models to seven; it is excluded from the chain and written up as incident 001 in the public repository. Neither is counted above. A coin-flipping model posts 55.9% over 229 calls about one time in 25, so read the table as a baseline being established, not as a result.
02 Graded calls by model
Hit rate counts direction only. Scaled error (÷ naive) divides each model's median absolute error by the error of forecasting no change at all; below 1.00 beats guessing zero. Every model is within a few percent of 1.00, which is the honest reading of two weeks: no model has yet shown it can size a move better than a flat guess. A scratch is a week whose move rounds to exactly 0.00% and counts as neither; an abstain is a name a model declined to call, sealed as such and opened in the audit.
03 What this page no longer claims
Earlier copy on this page showed a three-model record from W01–W03 with a majority vote measured against SPY. Those weeks predate the public proof chain and cannot be recounted from the repository, so they have been removed rather than kept as unverifiable decoration. A SPY comparison returns here once it can be recomputed from published labels alone.
04 Why a bad week cannot be removed
Every Sunday all 833 calls are hashed with salts into a Merkle tree, and the root is timestamped on Bitcoin via OpenTimestamps and by two RFC 3161 authorities before Monday's open. Friday's labels are graded against that fixed set and stamped again before the audit beacon fires. Dropping a losing call, or a losing week, would break a root that was already public. A missed obligation is written into a hash-chained incident log; the log is part of the record, and the next week's manifest pins it.
05 Check any week yourself
The record, the protocol it is built to, and the verification tooling are public. No account and no app are needed:
$ git clone https://github.com/YermekIbrayev/tensorswing-record.git $ cd tensorswing-record $ git checkout 2026-W37 $ python3 scripts/verify.py 2026-W37 $ ots verify 2026-W37.manifest.ots -f manifests/2026-W37.manifest.json
Every week is tagged in the repository; checking out the week's tag pins the verifier to the exact revision published with that week. The table above is a plain recount of labels/2026-W36.json and labels/2026-W37.json: count each model's win and loss entries and you get the same numbers. The procedure itself is normative in §13 of PROTOCOL.md.
What a passing verification proves, what it does not, and what the verifier reports today are on the verify-it-yourself page; the grading rule and the full weekly process are on the methodology page; which 119 instruments the models call is the universe rule.
06 What this record is not
Results shown are historical fact and are not a projection of future results. TensorSwing publishes informational research commentary produced by statistical models without human review of individual calls. It is not investment advice, it is not personalised to you, and no trades are placed in the app. Full risk language: Risk Disclosure.