Methodology
Last updated: September 2026
What this page is
A short, honest description of how Goalsforge reads a football fixture, runs the model, publishes the read, and audits it against the final score. It is not a research paper; it is a public description of the product so readers can decide how much weight to put on a read.
What the model reads
For each upcoming fixture the model is fed:
- Recent form. The last several matches for each side, weighted by venue and opposition quality. Raw form strings ("W-D-W-W-L") are deliberately treated as a noisy summary; the calibrated window priors the model uses are computed over the same window the upstream trends service uses (typically five matches).
- League base rates. The empirical 1X2 distribution, over-2.5 anchor, and BTTS baseline for the competition, taken from a full prior season. Deviations from base rate require evidence. Without it, the model collapses probabilities toward base rate.
- Standings and goal difference. Where the competition exposes them. Used to weight the strength of fixture context, not as a direct input.
- Head-to-head history. The recent head-to-head record between the two sides, scoped globally or to the competition depending on the call.
- Venue. Home vs away, and the kickoff window. The model treats late-Saturday and evening kickoffs differently for some competitions.
What the model does not read
- Real-time market prices. We do not ingest third-party market signals as an input. The model's probabilities are its own belief, not a smoothed market signal.
- Social media sentiment, expert narratives, or media storylines.
- In-play events before publish time. The read is generated and frozen before kickoff; live scores are a separate pipeline (the live read on the app) and do not rewrite the published read.
- Ad-hoc queries from readers. The model is not a chat. Every read is published and timestamped ahead of kickoff; the record is the answer.
How an analysis is produced
- Ingest. The fixture is identified and the inputs above are pulled from the local mirror.
- Feature build. A feature vector is assembled from form, base rate, standings, head-to-head, and venue.
- Model call. The model returns a single probability distribution over the canonical 1X2 outcome plus an estimated scoreline and a small set of secondary signals (likely scorers, half-time scenarios) when confidence allows.
- Validate. The response is checked for coherence: probabilities sum to 1, the modal scoreline is consistent with the 1X2 argmax, and the reasoning matches the probabilities. A response that fails validation is regenerated or, if the failure repeats, marked degraded.
- Publish. The read and the kickoff timestamp are written to the public record. The inputs that fed the call are captured at the same moment.
- Audit. After full-time the read is graded against the final score. The audit is a separate, idempotent pass. It does not touch the published read.
Calibration, not accuracy
The published figure on the/coverage/ index is calibration, not a raw accuracy count. A calibrated 65% read is one that lands right roughly 65% of the time across a large sample of similar fixtures. Not a read that matches 65% of matches. Calibration is the right metric for a probabilistic product. A raw match count rewards high-confidence favourites; calibration rewards being honest about uncertainty.
Publish and audit
Publishing an analysis means capturing its inputs and the call itself at the published kickoff time. The publish timestamp is auditable on every audited row. The audit is a separate pass after full-time: the read is graded against the final score and the result is added to the public record. The process that makes a read auditable is the same one that protects it from hindsight.
What this methodology is not
- It is not a guarantee. A read is a probability statement about a future event, not a promise. Football is structurally uncertain; calibration reduces that uncertainty, it does not remove it.
- It is not advice. Goalsforge does not tell readers what to do with a read. The product is the model and the public record, not a recommendation.
- It is not a guarantee of future accuracy. The published calibration describes how the model performed on past fixtures; it is not a forward-looking statement.
Read next
- The public record: every fixture, published and audited.
- How to read a probability-backed analysis: a reader's guide to interpreting a read.
- Why analyses are published before kickoff: the rationale for the timestamp.