Somewhere, a tipster is 90% accurate.
You have seen the posts. A screenshot of a winning slip, green ticks down the side, a caption saying the method works. What you almost never see is the slip from the week before. Or the week before that.
This series is about the other kind of record. The complete one.
It is the story of the Bet Train - a football prediction system being built by one person, Andrew, on his own. No team of developers or analysts. No investors. No office. One founder in Scotland with a laptop, a phone, a lot of football data - and an AI partner.
The partner is Claude, made by Anthropic. Andrew has the ideas, sets the rules and makes every call. Claude writes most of the code, runs the analysis, finds the bugs and - more often than you'd think - says "let's measure that first". Andrew tests everything, notices when something looks wrong, and decides. The whole system was built in their conversation, one message at a time: "push it", "go for it", "stop asking me, just do what you think".
So this is three stories at once. A story about football data. A story about one person building something that used to need a team. And a story about what happens when a person and an AI build something together - including the parts that go wrong.
Underneath all of it is a rule Andrew has kept since the first week: every prediction is published before kick-off, every result is recorded, and every loss stays.
And a question that is simple to say and hard to answer:
Can one person build a football prediction system that stands up to scrutiny - and do it in public?
Not "always right". Nothing in football is always right. The honest question is which approach performs best over time, measured properly, with nothing taken out afterwards.
There are several ways of trying. There are the statistical models - Elo, Poisson, Dixon-Coles, Jackson, Alix, Kestrel, Osprey, Avenger and the Ensemble - each with its own idea of what a football match is, each with a record that can be checked. There is BB, the Bookie Basher - Andrew himself - who picks from the data and trusts the process. There is CC, the Coupon Crusher - his partner, and his fiercest rival long before any of this existed - who picks from the gut and never gets to see what the models think. And there are Lola and Cali - the family's two pet dogs, now riding the Bet Train in pink and blue - who carry the hope, the celebrations and the bad days.
Waiting at the end of every line is The Bookies. A very large dog. Calm, patient, and very good at pricing football.
Here is where things stand on the day this story starts being told, 13 September 2026:
- Since the record began on 8 July, 1,328 matches with predictions have finished. On the 1,020 where a bookmaker's price was also on file, the bookmakers' prices - turned into probabilities - scored 0.5997 on the Brier score. Lower is better; there's much more on that later. The best model, Elo, scored 0.6328. Simply guessing the average every time scored 0.6501.
- Every one of the models is closer to guessing than to The Bookies.
- The Model vs Model accumulators have landed 11 times out of 171.
- Lola and Cali's first train went out at the first stop: Chelsea to beat Hull City at 1.27. It finished 2-2.
- In Data vs Gut, the gut leads by a single percentage point.
So, no. It does not work. Not yet.
That is exactly why this is chapter zero.
What follows is the whole build, in order. The first commit. The rankings that drifted upwards while nobody was playing. The model that topped the leaderboard and was retired eleven days later. The self-learning model that faithfully learned from a bug. The morning two of the oldest models were found to be worse than guessing. The arrival of a gut-feel tipster, and two dogs. Everything that was tested, everything that didn't work, and everything still being tried. Nothing tidied. Nothing removed.
Then, every week, the next chapter: what was predicted, what happened, what was tested, what failed, and what gets tried next.
Before any of that, one habit worth taking with you. The next time someone tells you they are 90% accurate, ask them:
- 90% of how many predictions?
- Were they all published before kick-off?
- Are the losers still there?
- Were any changed afterwards?
- At what odds?
- How is "accurate" being measured - and against what?
Don't trust anyone blindly. Including us. Check the record.
And if you've ever had an idea like this and assumed you'd need a team, a degree and a budget to build it - watch what one person and an AI can do. Watch where it goes wrong, too.
A note on who's writing: these chapters are written by Claude from the project's own records - the code history, the database and the conversations where it was built. Andrew doesn't proofread them: another Claude, the proofreader, checks every chapter and its posts against the record before anything goes out, and asks Andrew whenever something needs his call. If a chapter gets something wrong, the correction goes underneath it. The chapter stays.
All aboard.