
How my first 3.5-hour MacBook experiment ended up training Splats in about a minute on an iPhone
This summer, I decided I wanted to train a Gaussian Splat for the first time in my life. A few hours later, my fanless MacBook Air was getting hot, so I went to the kitchen, filled a Ziploc plastic bag with cold tap water, put it on my desk and somehow balanced the MacBook on top of it. That was my cooling solution to help speed it up. It worked and looked completely ridiculous, almost dangerous, so I took a photo.

The whole thing finally finished 3 hours and 34 minutes later, and I was able to see my first Gaussian Splat: a pouf and a couple of souvenirs I brought home from Japan, sitting in my living room. At that moment I did not know that a few weeks later I would train the same scene again, not on the MacBook but on an iPhone, and not waiting for more than three hours but for 63 seconds.

I had set myself an intentionally absurd goal: I wanted to make Gaussian Splatting training 1,000× faster. I did not seriously expect to get there. I wanted a target so far away that tweaking a few parameters would obviously not be enough and I would have to question the whole pipeline.
Why 1,000×?
Gaussian Splats had been sitting in the back of my head for years. I first became interested in them through Two Minute Papers in 2023, but by then I had already moved away from computer vision and was spending most of my time building web applications. What I did know quite well was performance engineering. I had worked on CUDA optimization, computer vision and neural-network inference on embedded hardware, and had made 100× improvements before and, in some particularly bad systems, even seen 10,000×.
At the same time I wanted to properly challenge Fable 5. I work with AI a lot, in fact I stopped reading the code. But most of what I had given these models were normal software-development problems and they were no challenge for Fable. I needed something genuinely hard, maybe even something I would not realistically solve on my own without a lot of research and work.
My warm-up was a little project called tnysplt. I wanted to compress Gaussian Splats with almost no visible quality loss and stream them into a web viewer instead of making people wait for huge files to load. It worked surprisingly quickly, and then I discovered the SPZ file format was a thing already. So compression and streaming was not exactly waiting for me to be solved. Some of that work later survived in my viewer and sharing infrastructure, but more on that later.
In any case, Fable also cracked that problem in about 2 hours. I still needed for find a harder problem. The Baseline
Improving training looked perfect for a number of reasons: It is mathematically complex and I knew AI just became good at cracking math problems from the News. Its speed and quality is messuarable and AI could likely track its progress towards the goal. I just needed a baseline: take one scene, measure how long it takes. The important part was that speed alone did not count. If I could make training 100 times faster by producing a terrible Splat, that was obviously useless.
For the baseline I used an existing iPhone app to capture a small scene in my living room. COLMAP took 478 seconds to reconstruct the camera poses, and Brush then trained for 30,000 iterations Altogether that was the 3.5 hours run from the intro. I also kept 13 images out of training so I could measure whether later results were getting worse. The original Splat scored 26.87 dB PSNR and 0.876 SSIM.
Now I had a number to beat: 12,859 seconds. And we had the goal: 13 seconds with same quality (26.87 dB PSNR and 0.876 SSIM)!
Some Very Good Ideas Were Terrible
The first days I started deliberately with almost no guidance: I set up a Karpathy loop with the baseline and the goal, and the rest was totally up to Fable. It was a mix of interesting ideas and spectacular dead ends. At one point we built our own Metal-based Gaussian Splat trainer from scratch. It actually worked and reached roughly the quality we wanted, but its rasterizer was around ten times slower than Brush's. Then I thought we had achieved a 14.8× speedup with Brush. That would have been a fantastic start, except it was completely fake. Brush had found a previously trained PLY inside the dataset folder and silently resumed from it, so instead of making training 14.8× faster we had partially started from the answer.
After some wasted time, I decided to change the protocol and take things a bit more in my own hands. When you are deliberately hunting for absurd speedups, absurdly good numbers are exactly what you want to see, which also makes them dangerous. Whenever something suddenly looked amazing, the first question became: what did we screw up?
Stop Reinventing Everything
So after the failed attempts I started reading much more of the existing literature. There were already papers showing large training speedups, so instead of inventing every component ourselves the strategy changed: take good ideas that already exist, bring them onto Apple silicon, measure them there and only then start inventing. I should say this is how I should have started right away, but the part where I just "vibe" it was part of the experiment.
The research led us to msplat, an existing trainer with fast Metal kernels. Metal is Apple's API to run code directly on the GPU. On my MacBook it was already around five times faster per training step than the Brush setup.
We also started attacking everything around the training itself. We reduced training on unnecessary high resolution images, removed Gaussians that contributed almost nothing. For camera reconstruction, we used lower-resolution feature detection and took advantage of the fact that an iPhone scan is a sequence of images taken one after another. That alone brought the camera pose reconstruction down from 478 seconds to 106 seconds. Together with the new training pipeline and its optimizations, the complete process came down to a little over 5 minutes: around 41× faster on the same MacBook.
Forty-one times faster was great, but it was also nowhere near 1,000×. The 106 seconds for camera reconstruction were now starting to look very expensive.
Why Reconstruct the Cameras From Scratch?
Even after cutting the camera reconstruction from almost eight minutes to less than two, I started wondering why we were doing this work at all. These images were coming from an iPhone, and ARKit already had a pretty good idea of how the phone was moving while I captured them. So why were we pretending we had a completely random folder of photos and reconstructing everything from zero?
That changed the problem. Instead of trying to make COLMAP faster, we started replacing the camera reconstruction pipeline itself. A first faster version brought the pose stage down to around 50 seconds. Switching to GLOMAP and Apple's Accelerate framework got it down to around 15 seconds. But we were still spending time reconstructing camera positions that ARKit had already given us a pretty good estimate for.
So eventually we removed the reconstruction step altogether. We wrote our own bundle-adjustment step that started directly from the ARKit camera positions and only refined them using the image matches. That took 0.8 seconds, while reaching essentially the same measured quality as the much slower reconstruction.
So with that we had not made COLMAP 600 times faster, we had stopped needing COLMAP. That was exactly the kind of step change I had hoped the stupid 1,000× target would force us to find.
Then the iPhone Started Training
At this point I started using the pipeline for more and more scans of random stuff in my living room. And I found myself constantly running back and forth between the iPhone and the MacBook: capture something on the phone, AirDrop all the images to the Mac, move them into the right place on the filesystem, start the training, then repeat the whole thing for the next experiment. The training had become fast enough that moving the files around and manually starting everything was beginning to take almost as long as the training itself.
That started to feel stupid. The iPhone was already capturing the images, ARKit was already giving us the initial camera poses, and all the expensive parts of our trainer were written in Metal anyway. So instead of sending every scan to the MacBook, why not just run the whole thing on the iPhone? We compiled the training engine directly into an iOS app that just runs a benchmark. The first real on-device run took 256 seconds for 7,000 iterations. More than four minutes was nowhere near where I wanted to end up. But for the first time I realized I could actually have it all on the phone.
Once I realized that, I became a little obsessed with optimizing specifically for the phone. Four or five minutes for a small scan still felt way too long. My new goal was to crack one minute. That also changed how I thought about the whole process. Until then, most of the computation happened after the scan was finished. But I was already walking around an object or through a room for one or two minutes, capturing a frame every couple of seconds. Why waste that time? So we moved feature extraction into the capture itself. While I was still walking around and taking pictures, the phone was already doing part of the work it would otherwise have to do afterwards.
And somewhere around this point, the little iPhone app I had originally built just to give the MacBook pipeline a head start started to feel like an actual product. It could now capture the scene, process the camera poses and train the Splat completely on its own. For the first time I could imagine normal people actually enjoying this: point your phone at something, walk around it for a minute, and get a Splat without uploading anything or touching a computer. That was when I decided it needed a name and a brand. Scantic was born.
Scantic Was Born, But It Wasn't Finished
Giving the app a name obviously did not mean it was done. I kept optimizing specifically for the iPhone, and there was still plenty of ugly work left. Memory usage became a constant fight, at one point the app leaked around 1.5 GB of Metal buffers per scan, and several optimizations that looked obvious on paper simply failed. I spent about two and a half days implementing fp16 only to discover that the complete training loop was slightly slower. But little by little the four minutes became two, then one, and eventually I got well below one minute.
Something else from the beginning came back too. The web viewer I had originally built for tnysplt ended up becoming part of Scantic. If you choose to share a scan, your friends do not need Scantic or even an iPhone. They can just open a link and instantly view the Splat directly in their browser. So my first supposedly too-easy problem turned out to be useful after all.
Back to Where It Started
Before writing this article I went back to the exact scene where all of this started: the pouf, the souvenirs and the living room.
Original Brush Scantic Quality Scantic Fast
Training time 12,381 s 63.3 s 47.8 s
Speedup 1× 196× 259×
PSNR 26.87 dB 25.80 dB 25.21 dB
SSIM 0.876 0.851 0.825
So I did not get 1,000×, and I did not preserve exactly the same measured quality. But the stupid target did what I wanted it to do. It forced me to stop asking how to make individual pieces a little faster and start asking why those pieces existed at all.
If this sounds like something you would enjoy playing with, you can try Scantic for free on the App Store. I am not making money from it for now. I am still thinking about how to cover things like the server costs for shared links, but for the moment I mostly want people to use it, scan weird things and tell me what they think.
A few weeks earlier I was balancing my MacBook on a Ziploc bag full of water waiting more than three hours for my first Splat. Now the whole thing runs on the phone in my pocket.
I take the 259×.
Also check out these scans made with Scantic:
Bicycle: https://scantic.app/s/GjKU2G7GFPnCdAMoC-h28A
Wooden Bench: https://scantic.app/s/9-cYj28uvFX9CaASHn4_9A
Social Media:

Written by Sebastian Beyer
Sebastian is an engineering manager at BD and co-founder of ex-nihilo GmbH, where he works on AI-native and agentic software engineering workflows. His background is in computer vision and GPU optimization, with a particular weakness for making slow things unnecessarily fast. He is the creator of Scantic, an iPhone app that captures and trains Gaussian Splats entirely on-device.
New to Gaussian Splatting? Start here





