Football Data · 20 September 2026

Building a model that knows when not to trust its data

tactica. found that having football data wasn't enough. The model also needed to understand whether evidence was valid, complete, timely and when not to trust it.

During preparation for tactica. v1.1, we found that having football data wasn't enough. The model also needed to know where it came from, when it existed, whether it was complete — and when not to use it.

We knew who the player probably was.

We refused to use that answer.

During an audit of the evidence underneath tactica., we encountered a fixture where our football-data provider returned 36 player records.

Thirty-five had usable identities.

One didn't.

The player was Ben Childs of Rotherham. In that particular provider response, his player ID was:

0

Elsewhere, we had evidence associating the player with a valid provider identity.

So there was an obvious solution.

Correct the bad record ourselves.

We didn't.

Because the moment tactica. substitutes an identity that wasn't actually present in the source record, something important changes.

We've moved from:

what the provider told us

to:

what we think the provider probably meant.

That might produce a cleaner dataset.

It doesn't produce more authoritative evidence.

So instead of repairing the record into something we wished we'd received, tactica. preserved the complete raw response, quarantined that individual row, admitted the other 35 validated players and labelled the resulting evidence:

VALID_PARTIAL

Thirty-five usable players.

One explicitly unusable player.

Not 36 supposedly perfect records.

That distinction captures a surprisingly large part of what tactica. v1.1 has actually been about.

Having data isn't the same as being entitled to trust it

It's easy to think about football-data infrastructure as a collection problem.

Do we have the fixtures?

Do we have the players?

Do we have the statistics?

Do we have enough history?

Those questions matter.

But preparation for tactica.'s public launch has increasingly forced us to ask harder ones.

Where did this evidence come from?

Does the source record actually say what we think it says?

Is the evidence complete?

When did it become available?

Was it available when the prediction was made?

If something is missing, does the model know it's missing?

And ultimately:

Is tactica. actually authorised to use it?

That has required us to audit the route from provider data through canonical football evidence and eventually towards the information available to Match Intelligence.

What we found wasn't one dramatic broken model.

It was more uncomfortable than that.

There were places where historical team and player evidence had not always reached the canonical observation layer we expected.

There were pagination and season-authority problems capable of leaving historical player evidence incomplete.

There was contextual evidence whose freshness needed to be explicit.

There were assumptions in our ingestion code that didn't always match the schemas our provider actually returned.

And there was a broader problem with time.

A model shouldn't merely know whether evidence exists.

It needs to know whether that evidence belonged to the world at the moment the prediction was made.

We went looking for 568 missing pieces of evidence

Once we understood the problem, we stopped treating historical recovery as an open-ended exercise.

We defined exactly what was missing.

Across the Premier League, Championship, League One and League Two, the recovery programme eventually froze a manifest of:

568 evidence requirements.

Then we worked through them.

Not by weakening validation until everything turned green.

By improving the acquisition and validation path until every requirement had an explicit, defensible state.

The final reconciliation was:

568 / 568 accounted for.

Of those:

567 were fully valid.

1 was VALID_PARTIAL.

That one was the fixture containing the provider's unusable player identity.

We could have made the headline prettier.

568 / 568 VALID

would certainly look cleaner.

It would also be false.

The point of the exercise wasn't to eliminate the appearance of uncertainty.

It was to represent the uncertainty correctly.

Partial means partial

This sounds like semantics until you think about what happens downstream.

Imagine a model asks:

Do I have player evidence for this fixture?

A crude system might answer:

Yes.

After all, we received the response.

Another might answer:

No.

After all, one row was defective.

Neither tells the full story.

The more useful answer is:

We have validated evidence for 35 of the 36 returned player records. One source row failed identity validation and has been excluded.

That's what VALID_PARTIAL represents.

It doesn't mean the evidence is good enough for every downstream use.

It doesn't mean the model automatically gets permission to use that fixture.

It doesn't mean we know how to repair the missing row.

It tells the next part of the system what evidence actually survived validation.

That is a much more useful thing to know.

The objective isn't to make uncertainty disappear.

It's to stop uncertainty disappearing without us noticing.

A prediction shouldn't know the future

Identity is one kind of authority.

Time is another.

Suppose tactica. generates Match Intelligence for a fixture on Saturday.

Then another match takes place on Wednesday.

Wednesday produces useful new evidence.

New team statistics.

New player minutes.

Potentially new availability information.

That evidence should absolutely be available for the next relevant prediction.

But it cannot be allowed to travel backwards in time and become part of what tactica. supposedly knew before Saturday.

That sounds obvious.

Enforcing it through a real football-data pipeline is less obvious.

If a system simply asks for the latest team profile, the latest player evidence or the latest contextual state, it can accidentally give a historical prediction knowledge that didn't exist when the prediction was made.

The prediction becomes better informed in retrospect.

The audit becomes less truthful.

And any attempt to evaluate what the model genuinely knew at the time becomes contaminated.

So one of the important directions in v1.1 is target-as-of authority.

Conceptually:

Prediction for Saturday → evidence available before Saturday.

Not:

Prediction for Saturday → whatever we know about that team today.

Historical evidence can remain factual while still being temporally invalid for a particular decision.

That distinction matters if we want tactica. to be genuinely auditable.

The record matters even when the journey isn't pretty

The 568-requirement recovery eventually consumed 757 provider HTTP calls.

That number is larger than the final evidence requirement.

There is a reason.

Some calls happened while we were discovering and correcting assumptions in the validation path.

We didn't subsequently rewrite the history so the finished process appeared to have worked perfectly from the beginning.

Those calls happened.

They remain part of the operational record.

Again, the objective isn't to create the neatest possible retrospective.

It's to preserve what actually happened.

That principle matters much more when we get to model Performance.

If tactica. is going to publish winners and losers honestly, preserve historical recommendations rather than rewrite them, and show what the model actually believed before a match, then the same attitude needs to exist much further down the stack.

Accountability can't begin at the final score.

Sometimes the safety system should stop you

We encountered the same principle somewhere completely different during the v1.1 work.

Before changing production authority structures, we wanted a verified database backup.

The production database had grown substantially.

Our existing backup safety check looked at the available working capacity and refused to continue.

That was inconvenient.

We wanted the backup.

We wanted to move forward with the work.

The easy response would have been to treat the guardrail as the problem.

We didn't.

We increased the available persistent storage, created a consistent backup, moved it to private storage and then downloaded that exact backup into an isolated non-production environment.

The checksums matched.

The database integrity check passed.

The restored schema matched.

Only then did we proceed with the additive authority work.

The interesting part isn't the backup technology.

It's that:

A safety check that stops a deployment isn't an inconvenience when it's correct.

It's part of the product.

The right response to a legitimate safety failure is not to weaken the safety mechanism until the operation succeeds.

It's to make the operation safe enough to pass.

“Do we have data?” is no longer a good enough question

This is probably the biggest change in how we're thinking about tactica.'s evidence.

The old question is simple:

Do we have some data for this match?

The better questions are harder:

What exactly is this evidence?
Where did it come from?
Was the source response valid?
Is it complete?
When was it available?
Is it still fresh?
Does it contradict anything else?
And is this particular part of tactica. authorised to use it?

Those questions create states that aren't always comfortable.

Valid.

Partial.

Unavailable.

Contradictory.

Stale.

Too late to belong to the historical prediction.

But those states are more useful than turning everything into yes or no.

We've already learned this lesson elsewhere.

Missing evidence isn't evidence of availability.

Something being calculated doesn't mean it affected a prediction.

Current information doesn't prove that the model possessed it when an earlier decision was made.

And now:

Knowing what a defective provider record probably meant doesn't make our inference part of the original evidence.

The same principle keeps appearing in different forms.

This doesn't prove tactica. will predict football better

That's an important limitation.

None of this demonstrates that tactica. will become more profitable.

It doesn't prove predictive accuracy has improved.

And 568 reconciled evidence requirements is not an accuracy statistic.

What this work does is improve our ability to establish what evidence exists, whether we trust it and what authority it should have.

Whether better-controlled evidence ultimately produces better football predictions is a separate question.

Performance has to answer that.

We shouldn't use infrastructure improvements as a substitute for that evidence.

v1.1 isn't finished yet

This work is part of the preparation for tactica. v1.1 and public launch.

At the time of writing, that programme is not complete.

Acquiring the missing evidence is not the same thing as proving every downstream observation, profile and model dependency is ready.

The remaining recovery and production-verification work still has to earn its own authority.

That matters because it would be very easy to turn a successful 568/568 acquisition result into:

Problem solved.

We don't think we've earned that sentence yet.

What we have earned is narrower.

We now have a much stronger contract for deciding what provider evidence is allowed through the door.

Predictions are easy to display. Trust is harder to engineer.

The visible part of a football model is the prediction.

A percentage.

An expected score.

A recommendation.

Those are relatively easy things to put on a screen.

The harder work happens underneath.

Can we account for the evidence behind the answer?

Can we distinguish what was known from what was inferred?

Can we preserve partial evidence without pretending it's complete?

Can we prevent today's information from quietly contaminating yesterday's reasoning?

Can the system refuse evidence when it cannot establish its authority?

And when something is uncertain, can it preserve that uncertainty instead of manufacturing certainty because a complete-looking dataset is more convenient?

That's what much of tactica. v1.1 has actually been about.

Not making the model look more complicated.

Making it more disciplined about what it is entitled to believe.

Because a football-intelligence system shouldn't just know how to produce an answer.

It should know when not to trust the data trying to give it one.