The 45 MB model we ship cut a pair of trainers out of a busy street scene and kept only the laces. Everything else in that photo, the soles, the mesh, the logo, came out as background. On the same photo a model a tenth of its size kept most of the shoe.
That was one of 31 test photos we ran through our background remover on 24 September 2026, all of it on the laptop itself: no photo went to a server, ours or anyone else’s. We wanted to know how good a cut-out you can get without uploading anything, and where it goes wrong: hair, fur, glass and product shots.
How we could score it
Judging cut-outs by eye is how most comparisons work, and it hides a lot. We built photos where the right answer is known. Wikimedia Commons has cut-outs that people made by hand and saved with a transparent background: portraits, cats and dogs, bottles, cameras, shoes. We placed 31 of them on four public-domain background photos (a crowded market, a forest path, a café kitchen, a living room), saved each as a JPG, and asked each model to find the subject again. The original cut-out’s transparency is the answer key.
The score is the average difference between the transparency the model produced and the true one, over every pixel: 0% is perfect, and 1% is roughly a thin wrong line around the subject. A photo where the model loses half the subject scores 20% or more. We ran everything in Chromium 153 on an Apple M5 Max laptop.
Two small models, and why we ship both
| Model | Size | People | Animals | Glass | Products | All | Time |
|---|---|---|---|---|---|---|---|
| IS-Net, main (Apache-2.0) | 44.6 MB | 4.2% | 2.5% | 0.7% | 3.0% | 2.6% | 0.10 s / 6.3 s |
| U²-Net small, second (Apache-2.0) | 4.4 MB | 3.5% | 1.8% | 1.0% | 1.7% | 1.9% | 0.03 s / 0.8 s |
| ormbg, people (Apache-2.0) | 84.0 MB | 2.0% | 4.7% | 1.2% | 4.7% | 3.6% | 0.06 s / – |
| MODNet, portraits (Apache-2.0) | 24.7 MB | 2.1% | 5.4% | 14.3% | 16.5% | 10.2% | 0.03 s / – |
| BiRefNet lite (MIT) | 109.2 MB | Ran out of memory |
Read only the last column and the small model wins: 1.9% against 2.6%. We ship the bigger one as the default anyway, and we think that’s the right call. It beat the small model on 20 of the 31 photos, including 9 of 10 products and 5 of 6 glass objects, and its typical photo is far cleaner: half the photos scored under 0.67%, against 1.28% for the small one. Its average is dragged up by a few photos it gets badly wrong: the trainers at 24%, a raccoon dog on the market background at 8%, and portraits on busy backgrounds, one at 6.7% because it kept part of the crowd behind the man.
The small model sees the photo at 320 pixels square, the big one at 1024. On an 18-megapixel portrait the difference shows at once: the small model’s outline around the hair is a soft blur several pixels wide and trims into the hair, where the big one follows it closely.
On people and pets the order flips. The small model did better on 3 of 5 people and 6 of 10 animals. So the tool offers the second model with one click under every result, and it only downloads if you ask. Had we always kept the better of the two for each photo, the error across all 31 would have been 1.5%.
The two specialists behaved as their authors say. ormbg, trained on people, was the best of all four on people at 2.0% and much worse on everything else. MODNet, built for portraits, cut out a camera and an iPhone as if they were background. BiRefNet, the model we most wanted to use because its weights carry the plainest licence of the lot, never ran: the lite version stopped with an out-of-memory error on its first photo, on both of the ways we tried to run it.
Hair, fur and glass
Hair is where every model loses most. We measured the error again in a thin band around the true outline, where the hair is, and it ran to 7.8% for the main model on people and 10.1% for the small one. You get a clean, slightly soft edge, not every flyaway strand. Fur was a little better on both.
Glass we can say less about than we hoped. The people who made our cut-outs cut the glass solid: a bottle is opaque in the answer key, so we measured the outline, not the see-through part. There both models did best, 0.7% and 1.0%. The one bottle whose maker had left the neck half-transparent came out solid from the main model. If you need a glass to stay see-through, you will have to paint that back yourself.
What we changed after measuring
One of our ideas made things worse, one helped a little and one helped a lot. We tried snapping the edge to the photo’s own detail with a guided filter, a standard trick for sharpening a matte. On these photos it raised the error at every setting we tried, from 2.8% to between 2.9% and 3.6%, so the tool doesn’t use it: the matte is scaled to the photo’s full size and applied to the untouched original.
We also re-estimate the colours of half-transparent edge pixels, to take out some of the old background showing through. Composited onto white, that cut the colour error around the edge by 3.9%. Real, but small.
The one that helped: both models leave a faint haze, about 1 part in 255, over the whole background, and stop just short of fully solid inside the subject. Stretching the matte so anything under 5% counts as clear and anything over 95% as solid cut the main model’s error from 3.0% to 2.6% and the small one’s from 2.3% to 1.9%. Harsher settings scored better still, but our answer keys are hand-cut with hard edges, so the score rewards hard edges. We stopped at 5% to keep soft hair soft.
What you download, and how long it takes
The main model is 44.6 MB. The original file is 170.4 MB; we store its weights as 8-bit numbers and turn them back into full precision when it loads, which made no measurable difference to the score (3.0% before and after). The engine that runs it adds 27.4 MB, so the first cut-out costs about 72 MB of download, once; after that it opens from your device. Both come from this site, not a third party.
On this laptop an 18-megapixel photo (5184 by 3456) took 0.9 seconds from drop to result using the graphics chip, and 6.4 seconds on the processor alone, which is what you get where the tool can’t use the graphics chip. The processor path runs on a single core here, because using several needs a page setup our site doesn’t have.
Once you have a clear cut-out, the file format decides whether the transparency survives: PNG and WebP keep it, JPG can’t. Our transparency guide has the details, and the saved file carries none of the original photo’s camera details or location.
Sources
- IS-Net (DIS project), Apache-2.0
- U²-Net, Apache-2.0
- rembg, the ONNX exports we start from (MIT)
- ormbg on Hugging Face, Apache-2.0
- MODNet, Apache-2.0
- BiRefNet, MIT
- ONNX Runtime Web, MIT
- Forte and Pitié, Approximate Fast Foreground Colour Estimation (the edge colour fix)
- Wikimedia Commons: Merrell Vapor Glove 3 shoes (cut-out)
- Wikimedia Commons: Coca Cola glass bottle transparent (cut-out)
- Wikimedia Commons: Nyctereutes procyonoides (raccoon dog cut-out)
- Wikimedia Commons: John Thune portrait (no background)
- Wikimedia Commons: Busy street market (background)
- Wikimedia Commons: Twisting path in an autumn forest (background)
- Wikimedia Commons: Interior of Kitchen Cafe (background)
- Wikimedia Commons: living room, Edifício Zaher (background)
- Wikimedia Commons: Brunette woman portrait (the 18-megapixel timing photo)