Tim Shyu 中文EN日本語

How Much Data Actually Counts as "Big Data"?

Taiwan's mobile-payment wars have flared up again: one provider is reportedly planning to spend NT$500 million on a subsidy blitz, and rivals aren't backing down either. Few businesses willingly bleed money, so why are payment providers so determined to fight to the death over this? Ask them, and the pitch is usually the same: payments are a high-frequency, must-use behavior that generates enormous amounts of user data, and user data will matter enormously down the road, so we can't afford to fall behind now. In short, they believe that a mountain of user data will let them understand users deeply enough to seize whatever opportunity comes next.

Which raises an interesting question: how much user data actually counts as "a lot"? And once you've amassed all this big data, what exactly can you do with it?

There's a great illustrative anecdote from TSMC president C.C. Wei. Wei tells the story that after the auto chip shortage hit, carmaker executives who'd never once reached out to him before started calling him up like old friends. He'd ask how many wafers they wanted to order. "Twenty-five," they'd say. Wei was stunned — TSMC needs an order of at least 25,000 wafers just to start a production run. No wonder you can't get support, he told them. A modern car needs a fair number of chips, sure, but today's vehicles use somewhere around 120 chips each — older models use fewer still — so however much the carmakers stamped their feet, there was nothing to be done.

What You're Trying to Do Determines How Much Data You Actually Need

This is a classic case of every industry speaking its own language. To a carmaker, an order of a hundred thousand or a million chips already sounds like a massive procurement number — big enough that suppliers show up to bid competitively for it. But to TSMC, that same number isn't even enough to justify a production run. Whether a number counts as "large" or "small" depends entirely on what you're trying to do with it. Big data's bigness, and small data's smallness, are both entirely relative.

Here's a more personal example: I often have clients tell me they've already spent several million, even ten million-plus New Taiwan dollars on digital advertising, and now they want to build some sort of powerful data application on top of it — enough to take on Google and Facebook. What's genuinely hard to explain is this: to the client, ten million-plus NT dollars in ad spend is a number big enough to lose sleep over. But Google alone pulls in over NT$6 trillion in global ad revenue a year, and Taiwan's entire digital ad market tops NT$54 billion annually — your ten million is roughly 1/5,400th of that. For reference, a rare disease like Marfan syndrome is classified as "rare" because its incidence is roughly 1 in 5,000. So building some sort of powerful data-driven application on a budget that small is, unfortunately, a tall order.

Most people badly underestimate just how much data giants like Google and Facebook actually hold. Just take stock of your own daily habits and you'll get a sense of it: do you open Google Maps? Do you check Instagram or Facebook at all? Is your phone running Android? Trying to out-data these giants is a lot like trying to beat the house at a casino — an exhausting, thankless endeavor.

Doing Something Useful Doesn't Require an Astronomical Amount of Data

Telling a client all that outright would obviously go over badly, so what's the actual advice? The good news: most companies don't need that scale of data to pull off a genuine digital transformation, because they typically have no shortage of small, purpose-specific data close at hand. Take the USPS: it knows very little about who its customers are, but there's one thing it knows extremely well — handwriting. The USPS's handwriting-recognition system can now correctly read 98% of mail labels, leaving the remaining 2% for human specialists to sort out. That system works so well precisely because the USPS didn't try to boil the ocean; it correctly identified that improving postal efficiency really only required handwriting recognition — and that kind of data is everywhere at the USPS, sitting on virtually every single piece of mail. It's essentially big data that's already lying around for the taking. If the USPS instead wanted to build an accurate profile of the people sending letters, it would have almost no relevant data to work with, the project would be a struggle, and frankly it isn't the USPS's core business pain point anyway. The right approach is always to start from the data you already have, and work backward to figure out what kind of "big data" system you actually need.

Another example: Vigilant Solutions, a data company that builds automated license-plate recognition systems for U.S. police departments. How much data does that actually take? The company claims to hold at least 5 billion license-plate records (that's records, not distinct plates — no typo there), plus another 1.5 billion data points contributed by various U.S. law-enforcement agencies. And all that data is aimed at exactly one thing: license-plate recognition. Both cases prove the same point — you don't need an astronomical volume of data to do something genuinely useful. Whether it's handwriting samples or license-plate records, these narrowly targeted systems still perform extremely well. At a moment when most algorithmic applications still haven't gone mainstream, improving the bulk of your operational efficiency simply doesn't require an enormous data set.

So, back to the original question: how much data is actually enough for the payments industry? I'd say it depends entirely on what the people running the show are trying to achieve. If the goal is a modest one, like improving member services, then what they have is plenty. But if the ambition is to build a monopolistic gateway over the long run — where the advertising lives with them, the data applications live with them, everything lives with them — then I don't think their data volume comes close to being enough. Even a behemoth like LINE Taiwan can only barely scrape by at that scale, let alone truly rival the giants that control the operating system layer itself. Data collection, after all, isn't like opening physical storefronts — the global platform giants can plant themselves directly inside your phone's core infrastructure whenever they please. And bundling all your data applications entirely in-house means that no amount of subsidy spending will be enough in the long run. A better approach might be for payment providers to support a more open data ecosystem instead — but that's probably a conversation that has to wait until the current subsidy war finally burns itself out.

The Tim Shyu Letter

First-hand notes on AI agents, marketing tech, and the content industry — straight to your inbox.

Subscriptions open when the site goes live. Hold tight.


← All articles