Bad Data Is Costing Sports Betting Operators · SportsStack
Articles

Bad Data Is Costing You Money: Building the Unified Data Platform for Sports Betting & Prediction Markets

Every operator rebuilds the same data plumbing. Settlement is where the cracks turn into real money. From my talk at SBC Summit in Lisbon, why sports data needs a Unified Data Platform and what it takes to build one.

A

Amar Rao

Founder & CEO, SportsStack

Amar Rao speaking at SBC Summit in Lisbon
SBC Summit, Lisbon. Photo supplied by Amar Rao.

The same problem, over and over

I spoke at SBC Summit in Lisbon about a problem I have spent 13 years working on: how do you know you can trust the sports data powering your product?

I have worked on this at numberFire, FanDuel, and Fanatics Sportsbook. Different companies, different products, the same plumbing. Integrate the providers. Map the IDs. Clean up the formats. Deal with the edge cases. Keep it all running while the next season, league, or market launches.

That is why I started SportsStack. Every operator needs this layer, but it should not be something every operator has to rebuild.

Amar beside the 13 Years in Sports Data slide
The same data problems across numberFire, FanDuel, and Fanatics Sportsbook.

Sports data is fragmented

A sports product can depend on separate providers for stats, odds, content, images, news, and projections. Each has its own IDs, schemas, formats, and quirks. The same player can have a different identifier in every feed. It is your job to turn all of those records into one player in your product.

Adding a provider means another integration. Switching one means doing a lot of that work again. And the work does not end when the feed goes live. Providers change their formats, coverage has gaps, and exceptions show up in production.

Every operator is paying engineers to solve some version of the same problem. That is expensive, and it takes time away from building the things that make the product different.

Redundancy is expensive. No redundancy costs even more.

A second feed has a cost. You have another contract, another integration, and another set of mappings to maintain. It is easy to look at that spend and ask whether you really need it.

We once had a provider go down for three weeks during the NFL season. That changes how you think about the cost of a backup.

The point is not to buy every feed. It is to build a system where you can add or switch providers without rebuilding the product around them. Redundancy only helps if your infrastructure can use it.

Your customers do not blame the provider

When a score is wrong or a bet is misgraded, the customer does not care which upstream feed sent the bad number. They blame you. They think your product sucks.

A data issue becomes a product issue the moment a customer sees it. And when that data is used to settle a market, it becomes a money issue.

Amar presenting Bad Data Is Costing You Money at SBC Summit
Bad Data Is Costing You Money: the talk behind this post.

Settlement is where data turns into money

Say you have a market on LeBron James over or under 3.5 assists in the first quarter. The quarter ends and one provider reports three assists. Do you settle?

Now imagine three providers agree on three, but an official correction changes it to four five minutes later. Agreement did not make the original number right. If you already paid out, you have a problem.

Or the feeds report three, four, and three. Two out of three is a majority, but that does not automatically make three the truth. Which provider do you trust for that stat? How long do you wait? When should you hold?

This is why I say settlement is an art, not a science. There is judgment in deciding when a number is reliable enough to pay on. You need rules for the provider, the sport, and the market, tuned to your risk tolerance.

Customers want fast settlement. Operators need correct settlement. Getting that balance right across an entire catalogue is a much bigger job than checking whether a feed says final.

Settlement example (illustrative): three assists reported, corrected to four after five minutes, then feeds disagree three, four, three
Illustrative example, not a real game.

The game was final. Until it wasn't.

I used the Western Michigan-Michigan game as a real example in the talk. Kalshi confirmed that it prematurely settled the market as a Western Michigan win, then reversed the outcome and corrected the payouts. NBC Sports reported $18.6 million in trading volume on the game. That was volume, not a reported loss.

That is the risk in treating a final flag as a payout instruction. A result can still be under review. Your settlement process has to account for that and follow the applicable official result and market rules.

This was not a SportsStack customer case. It is a public example of the problem operators face: a fast answer can still be the wrong answer.

What doing this well actually takes

You need a platform that abstracts away the providers, so adding or switching one is manageable. You need operational tooling to catch gaps and wrong values before customers find them. And you need a team that owns the data, maintains the integrations, and handles the exceptions.

For settlement, you also need more than one source and rules that reflect how each source behaves for each market. Some stats are more likely to change. Some sources are more useful for one sport than another. A conflict should be a reason to check, not a reason to blindly pick the majority.

None of this is cheap. It is also hard to fund because sports never stop for a rebuild. There is always another tentpole event, another launch, another deadline. Systems that work well enough keep getting patched, even when everyone knows they need more attention.

The Unified Data Platform we are building

SportsStack is the unified data platform for sports. We focus on three parts of that infrastructure.

The Mappings API connects provider-specific IDs to one identity for each player, team, and event. Your product should not have to figure out whether two feeds are describing the same LeBron James.

The Unified API puts the providers behind one interface and a consistent schema. Scores, stats, odds, projections, content, and news should not require a different downstream integration every time you add a source.

The Settlement API cross-validates results, identifies conflicts, and gives you recommendations on whether to hold or settle. The rules reflect your risk profile: how long to wait, which sources to trust, and which stats need more caution. Those recommendations fit into your existing settlement process.

These are modular products. If mappings are the problem you need to solve, start there. If settlement is where the risk sits, focus there. The goal is to give you better infrastructure without asking you to replace your entire stack.

Four sports data provider feeds enter SportsStack Unified API to give a product one consistent interface
One endpoint and a consistent schema across sports and data types.

Build your product, not the plumbing

I have spent much of my career trying to take systems that work and make them great. The lesson is that clean, trusted data has to come first. Everything built on top of it depends on that foundation.

You still own the customer experience and the decisions your business makes. But mapping IDs, normalizing providers, and checking whether the data is safe to settle on should not be projects every team has to start from scratch.

If you are building a sportsbook, prediction market, fantasy product, or anything else that depends on sports data, I would like to hear where this breaks for you. Reach out at amar@sportsstack.io.

SBC Summit sports data settlement misgrades unified API mappings prediction markets unified data platform

Talk to us about your data

One integration, universal IDs, any provider, and confidence in every market before you settle it.