Our model was profitable. We audited it anyway.

tactica. returned +4.96% across its first meaningful observational Performance sample. Instead of treating that as proof the system was healthy, we audited the intelligence underneath it.

Share
Our model was profitable. We audited it anyway.
A positive result can put a model in the spotlight. The audit starts with what the result cannot tell you.

tactica. returned +4.96% across its first meaningful observational Performance sample. That looked encouraging. It wasn't enough to tell us whether the intelligence underneath was healthy.

If a football prediction model is making money, why would you go looking for reasons not to trust it?

That was the question facing us.

Across 262 settled model positions covering 93 fixtures between 31 July and 24 August 2026, tactica. recorded £325.18 profit from £6,550 of fixed model stakes.

An ROI of +4.96%.

These weren't customer bets or somebody's personal betting record. They were model recommendations recorded specifically so we could measure tactica.'s Performance against actual outcomes.

The result was encouraging.

It would also have been very easy to misuse.

We could have taken the number, put it on a graphic and told everyone the model was working.

Instead, we audited it.

Profit tells you what happened. Not necessarily why.

A profitable period is an outcome.

It doesn't tell you whether the probabilities underneath it are well calibrated. It doesn't tell you whether one competition is carrying everything else. It doesn't tell you whether the model's explanations accurately describe what it calculated.

It doesn't even tell you whether the profitable parts of the model are the healthy parts.

So our audit asked a harder question than did tactica. make money?

We examined the chain underneath it:

Football evidence → calculation → probability → recommendation → explanation → outcome → accountability.

Could we trust each part?

That distinction became important very quickly.

Some of the model held up

There was good news.

Across 241 simulations, 1,928 Result-market probability transformations passed the mathematical identities we tested.

The way tactica. constructed its home, draw and away probabilities behaved coherently. Its Draw No Bet conditioning and Double Chance calculations also behaved mathematically as expected.

That's useful evidence.

But even here, there is a line we won't cross.

It doesn't prove tactica. is perfectly calibrated.

It doesn't prove the probabilities are objectively correct.

And it certainly doesn't prove tactica. is better than the bookmakers.

The maths behaving as designed and the model being predictively superior are two very different claims.

We have evidence for the first.

We don't yet have evidence for the second.

Then we found something we didn't like

One of the audit findings concerned the explanations tactica. stored about its own predictions.

We found 418 cases where player-quality evidence was missing on one side of the calculation.

In 415 of them, the explanation recorded a non-zero finishing contribution.

There was a problem.

That finishing contribution had not actually changed the final expected-goals prediction.

The calculation had produced the number, but because the necessary quality evidence was missing, it hadn't been applied to the final result.

The explanation nevertheless carried it forward as though it had contributed.

That's an important distinction.

We did not discover 415 wrong predictions.

We discovered 415 cases where the causal explanation of a prediction said more than the underlying calculation justified.

For a product whose entire proposition depends on users being able to inspect why it thinks something, that's not a cosmetic problem.

It's an authority problem.

We could have made the explanation true by changing the model

There was an obvious way to make the discrepancy disappear.

Add the finishing contribution into the expected-goals calculation.

Then the explanation and prediction would agree.

We refused to do that.

The audit had proved that the explanation was wrong.

It had not proved that the model should use the missing contribution.

Those are completely different conclusions.

Changing the prediction simply to make an existing explanation true would have turned an explanation defect into an unvalidated model change.

So we chose the opposite approach.

We built and verified a correction to the explanation contract while leaving the underlying predictions unchanged.

Across 313 pre/post fixture comparisons, the corrected implementation produced identical numerical outputs for expected goals, result probabilities, BTTS, count projections and market projections.

We changed what tactica. was claiming about its reasoning.

We didn't change what it predicted.

That distinction matters to us.

The winners weren't automatically healthy

The same discipline applied to Performance.

BTTS was the standout market in the sample:

36 positions. +£195.75. +21.75% ROI.

Sounds impressive.

Its estimated 95% fixture-bootstrap ROI interval ran from -8.84% to +49.97%.

In other words, the uncertainty remained substantial.

So we didn't declare BTTS tactica.'s winning market.

We didn't increase its recommendation frequency because it had performed well.

We kept watching.

Corners returned +4.93%.

Bookings returned +8.11%.

Yet the audit found that parts of the product language around those markets claimed richer football intelligence than the underlying evidence justified.

Positive returns didn't override that problem. Corners remained paused, while existing restrictions on some Bookings recommendations remained in place.

A profitable market can still have an intelligence problem.

The loser wasn't automatically broken either

Draw No Bet went the other way.

Across 50 positions it lost £62.43, an ROI of -4.99%.

We didn't disable it because it lost money.

The uncertainty around that result was wide, and the audit separately uncovered a more interesting problem: tactica. could compare a Draw No Bet probability with an outright-win probability even though the two describe different outcomes.

One includes protection against a draw. The other doesn't.

The maths behind the conditional probability could be coherent while the way we qualified the resulting recommendation still needed investigation.

The loss didn't prove Draw No Bet was broken.

The evidence around how we were using it earned further investigation.

Again, those are different things.

Then there was League Two

This was perhaps the easiest result of all to overinterpret.

League Two returned:

63 positions.
48 wins.
11 losses.
4 voids.
+£298.53.
+18.95% ROI.

Its fixture-bootstrap interval was also the only competition interval in this analysis that sat wholly above zero.

You can probably imagine the headline.

tactica. has found a League Two edge.

We don't think the evidence earns that sentence.

Those 63 positions came from just 24 fixtures across two matchdays: 15 August and 22 August.

Three fixtures alone contributed £153.20 of the profit.

Further analysis suggested that market mix by itself didn't explain League Two's apparent outperformance.

Interesting.

But we still couldn't establish whether we were seeing durable league-specific skill, an underlying tactical reason, better calibration, or simply an unusually strong short run.

So we didn't create a League Two coefficient.

We didn't lower its recommendation thresholds.

We didn't build a special League Two model.

And we didn't advertise a proven League Two edge.

League Two had earned our attention. It hadn't earned our intervention.

The evidence itself wasn't perfect

There was another uncomfortable part of the audit.

Our historical Performance record wasn't immaculate.

We found two settlement records that couldn't safely remain in the analytical sample.

One was demonstrably invalid.

For the other, conflicting evidence existed and the original settlement-time authority could no longer be recovered.

We excluded both.

That means even the 262-position corpus we're discussing here needs the correct description.

It is an observational Performance sample.

It isn't an independently certified historical betting record.

And the period itself was short: less than a month of early-season football.

Every market-level uncertainty interval in the audit crossed zero. Calibration analysis was possible on only a subset of compatible observations and came with timing and matching limitations.

Those aren't footnotes we want to hide.

They're part of the result.

Unknown isn't Available

Performance wasn't the only place where the audit forced us to confront uncertainty.

We also found cases where missing injury evidence could effectively become Available, and missing injury coverage could become an injury count of zero.

Neither conclusion was justified.

No evidence of an injury is not the same thing as evidence that a player is available.

So the contract became simple:

Unknown is not Available.

Missing injury evidence is not zero injuries.

That sounds obvious when written down.

Systems don't always behave according to what sounds obvious.

That's why we audit them.

So, was +4.96% good?

Yes.

We would rather the first meaningful observational sample had returned +4.96% than -4.96%.

But that's about as far as the evidence lets us go.

The same audit contained all of these things at once:

A profitable overall result.

A highly profitable BTTS sample that remained uncertain.

A losing Draw No Bet sample that didn't prove the underlying maths was broken.

A striking League Two result that wasn't strong enough to justify league-specific tuning.

Probability mechanics that largely behaved correctly.

Explanations that sometimes claimed causality they hadn't earned.

Markets making money while parts of their product authority remained questionable.

And missing evidence being represented with more certainty than it deserved.

There isn't a clean marketing slogan hiding inside that evidence.

Good.

We're not looking for one.

The rule we're carrying forward

We came out of this audit with three sentences that now govern how we think about model change:

Profit does not prove health.

Loss does not prove failure.

Evidence earns change.

That last sentence is the important one.

An audit shouldn't exist to find reasons to change a model.

It should establish whether change has been earned.

Sometimes that means fixing something.

Sometimes it means investigating further.

And sometimes it means looking at a result as attractive as +18.95% and deliberately doing nothing.

That's the standard we're trying to build tactica. around.

It's also why this Journal exists.

There will be winning results here.

There will be losing ones.

There will be things we build that work, things we build that don't, and occasions where the evidence tells us the most responsible decision is to leave something alone.

If we're going to ask people to trust tactica.'s intelligence, we think they should be allowed to see that part too.