TestFlight beta iPhone and iPad · iOS 18 or later

The AI benchmark that makes your iPhone sweat.

Same models, same prompts, measured the same way, so the only thing left to compare is the chip. RapidInference runs four real AI workloads on your iPhone, with nothing sent to a server, and gives it one score you can put next to anyone’s.

A20 Pro · two Neural Engines 86 tok/s Peak 2.4 GB No network needed

Home, after a full run. Tap Run full benchmark to watch it again.

01 / Live replayGemma 4 E2B · MLX Swift · GPU

This is 86 tokens a second.

Gemma 4 E2B answering the first question of the benchmark on an iPhone 18 Pro, replayed at the speed the phone measured. Eleven questions in, the chip is hot and the same model manages about 40. The score counts both.

GSM8K · problem 1 of 20 greedy · ≤512 tokens

Janet’s ducks lay 16 eggs per day. She eats three for breakfast every morning and bakes muffins for her friends every day with four. She sells the remainder at the farmers’ market daily for $2 per fresh duck egg. How much in dollars does she make every day at the farmers’ market?

First token
–
Tokens
0
Elapsed
000 s
Decode
– tok/s

Pace from the reference run on 22 September 2026. The wording of the reply is ours; the speed is the phone’s.

02 / The suiteANE → Inference → Training → Image

Four tests, in this order, for a reason.

The Neural Engine goes first. After a few minutes of GPU work the whole chip throttles, and measured last, the second Neural Engine’s advantage simply disappears. Then the language model, the fine-tune and the images, all on the GPU through MLX.

01 Neural Engine Runs first

Upscales a photo, tile by tile, on every engine the chip has.

EDSR ×2 super-resolution on 25 overlapping tiles of a harbor photo, swept across one, two and four lanes on the Neural Engine, then the GPU and the CPU.

145tiles/s on two engines
Watch the lanes
Model
EDSR r16f64, fp16 · 3 MB, built in
Runs on
Core AI: Neural Engine, GPU, CPU · iOS 27
Reference
145 tiles/s · 10.8 ms per request · +1.71 dB
02 Inference

Solves twenty grade-school math problems, back to back.

The first 20 problems of GSM8K, answered greedily, up to 512 tokens each. Speed is timed on the first five and the last five, so the score sees how much a warm chip slows down.

16 of 20 answered correctly on the reference phone, the same 16 as on a Mac.

Model
Gemma 4 E2B, QAT int2/int4 · 2.2 GB
Runs on
MLX Swift · GPU
Reference
86 tok/s decode · ~100 ms to first token
03 Training

Fine-tunes a language model, right there on the phone.

150 steps of QLoRA on 500 rows of WikiSQL, teaching Qwen3 to turn questions into SQL. It asks three everyday questions before and after, so you can see what it learned.

loss 3.25150 steps1.24

QWhat is the capital of Japan?

ASELECT Capital FROM countries WHERE Country = 'Japan'

Model
Qwen3 0.6B, 4-bit · 351 MB
Runs on
MLX Swift · GPU
Reference
1.26 steps/s · 150 steps in 119 s
04 Image

Paints three pictures from a prompt, in four steps each.

Text to image at 512 × 512 with a ternary FLUX.2 [klein]. The text encoder, transformer and decoder load one at a time, so the whole thing fits in about 3 GB.

  1. seed 7a meticulously pruned bonsai tree on a stone ledge, soft studio lighting, photorealistic
  2. seed 42a red fox curled up asleep in fresh snow, golden hour, shallow depth of field
  3. seed 1234a futuristic city skyline at night with neon signs reflected in rain puddles
Model
Bonsai Image, FLUX.2 [klein] 4B · 3.8 GB
Runs on
MLX Swift · GPU
Reference
16.6 s per image · 2.76 GB peak
03 / Neural EngineEDSR ×2 · Core AI

Measured on iPhone 18 Pro, 22 September 2026. Slowed down 10× so you can watch it happen.

One Neural Engine, or two?

The A20 Pro has two. To find out, RapidInference cuts a 1024 × 1024 harbor photo into 25 overlapping tiles and upscales each one ×2 with EDSR, sending them down one, two or four lanes at once.

A single engine takes one request at a time, so extra lanes just queue up. With two, the second lane runs alongside the first: 145 tiles a second instead of 85.

tiles per second · pick one to run it

04 / The scoreScore version 1.0

5,000 is an iPhone 18 Pro. The rest is arithmetic.

Like Geekbench, points are anchored to a real phone. Twice as fast scores twice as much, and there’s no ceiling. Only speed earns points: a wrong answer can take them away, never add them. Memory and heat sit beside the score as badges, not inside it.

Your phone 1.00× the speed
5,000

points = 5,000 × your speed ÷ the reference · bars fill at 6,250

Inference

  • Decode speed, first 5 prompts40
  • Decode speed, last 5 prompts35
  • Median time to first token25

Lowered by wrong GSM8K answers. Badges: peak memory, and how much speed holds.

Training

  • Steps per second60
  • Total time, load and evals included40

Lowered by a worse validation loss. Badge: peak memory.

Image

  • Images per minute60
  • Denoising steps per second40

Badge: peak memory.

Neural Engine

  • Tiles per second, best lane count60
  • Single-request latency40

Lowered if the upscale gains less sharpness over plain resizing.

Each test is the weighted geometric mean of its speed metrics; the overall score is the geometric mean of the four. Quality can lower a score to half, never raise it. Reports keep raw numbers, so old runs are rescored when the reference moves.

05 / DetailsThe parts you might not notice

We sweated the small stuff too.

Downloads that don’t need you.

The models download in the background, even with the phone locked. A Live Activity follows them on the Lock Screen and in the Dynamic Island, in 5% steps, and tells you when they’re ready.

0 B

of your prompts, pictures and reports leave the phone. Scores go to Game Center only if you’re signed in.

64 GB

of models, fetched once, resumed if the Wi-Fi drops, and kept out of your iCloud backup.

It waits for the phone to cool.

A phone that’s already warm waits up to five minutes before the suite starts, so no run begins throttled. In a hurry? There’s a Skip button.

Every run, kept.

Reports save the raw numbers and the images, not just the score.

  • ri1.overall
  • ri1.inference
  • ri1.training
  • ri1.image
  • ri1.ane

Ranked by phone, too.

Five Game Center boards. Every score carries its chip, so phones get ranked, not just players.

Simulator · not submitted

No asterisks.

Simulator runs are never submitted. A run you leave halfway isn’t scored. Fixed prompts, fixed seeds, fixed order, every time.

06 / BetaTestFlight

Put your iPhone on the bench.

RapidInference is in beta on TestFlight. Leave your email and we’ll send you the link.

  • iOS 18 or later
  • Neural Engine test needs iOS 27
  • About 6.4 GB of models
  • Wi-Fi recommended