A smartphone processor now has more cache memory than most laptop CPUs on the market. That sentence alone should reframe how people think about who’s actually pushing chip design forward in 2026 — and it isn’t the company most people assume.
The Chip
Xiaomi unveiled its Xring O3 processor this week, built on TSMC’s 3nm process with roughly 24 billion transistors, according to multiple outlets including HotHardware, Digital Trends, and Android Headlines. The chip uses a 10-core, all-big-core design split across three clusters: two C1-Ultra prime cores clocked up to 4.35GHz, four C1-Premium cores up to 3.68GHz, and four C1-Pro cores up to 3.15GHz, Digital Trends reported.

Xiaomi claims Geekbench 6.5 scores of 3,945 single-core and 15,221 multi-core, improvements the company says represent gains of 31% and 61% respectively over its previous Xring O1 chip, according to Digital Trends.
Tech outlet NokiaPowerUser reported the chip carries a unified 60MB total cache, and separate reporting from vgtimes.com noted the die includes a new 16MB system-level cache layer, with total die size growing from 109 to 133 square millimeters compared to its predecessor.
Software performance researcher Daniel Lemire, who ranks among the top 2% of scientists globally according to a 2025 Stanford/Elsevier analysis, wrote in a widely shared post on X that the chip’s architecture looks legitimately impressive, while cautioning that the claimed benchmark numbers also carry “somewhat BS” energy until independently verified. Lemire pointed specifically to the chip’s C1-Ultra cores, noting they support SME2 (Scalable Matrix Extension 2) for AI and matrix acceleration and SVE2 for data-parallel SIMD workloads, and that the design carries 21 execution ports, six of which handle 128-bit SIMD operations — more execution ports, he noted, than found on typical Intel or AMD desktop processors.
The chip is set to debut in the Xiaomi 18 Fold and Pad 9 Pro Max, with NokiaMob and the Free Press Journal both citing an AnTuTu v11 benchmark score above 5.22 million, which several outlets described as the first mobile chip to break the five-million-point mark on that benchmark. HotHardware reported initial production shipments are targeted at 200,000 to 300,000 units, with independent lab testing still needed to confirm how the design handles sustained thermal loads.
Why the Cache Number Matters More Than the Benchmark Score
Lemire’s framing is the one worth taking seriously here, and it’s not really about whether Xiaomi beat Apple’s A19 Pro on a synthetic benchmark chart. It’s about where the transistor budget went. Sixty megabytes of cache on a phone chip, more than most laptop-class Intel or AMD parts carry, signals that Xiaomi isn’t just chasing peak clock speed, it’s building a chip that keeps far more data close to the execution units, which is exactly the kind of design choice that pays off in real, sustained multithreaded workloads rather than benchmark bursts. Combined with a 21-port execution engine, this isn’t a company copying ARM’s reference design with a marketing coat of paint. It’s a genuinely wide, cache-heavy architecture aimed at AI and parallel workloads specifically.
The bigger story is what this says about the smartphone chip race generally. Xiaomi, MediaTek, and Qualcomm are now routinely trading blows with architectural decisions that used to be the exclusive domain of desktop and server silicon, matrix extensions, huge unified caches, wide execution ports, while companies like Intel remain comparatively conservative on cache allocation in mainstream laptop parts. If Lemire’s skepticism about the benchmark numbers proves right and the real-world gains are smaller than claimed, that’s a normal part of every chip launch cycle. But the architectural direction he’s flagging, more cache, wider execution, dedicated matrix units, isn’t a one-off marketing gimmick. It’s where the entire mobile chip industry is now headed, and PC chipmakers watching from the sidelines should probably be paying closer attention than they currently seem to be.
