local ai mimo v26 qwen 9b
ShareByLight

I just launched ShareByLight—transfer files with light only. No Wi-Fi, no Bluetooth, no cables. One device shows a stream, the other reads it with the camera. Free on iPhone, iPad, and Mac.

This is the third video already in the series I started this month. I am trying to share my experiments in running relatively small – but powerful – models on an entry level machine, my 16GB M1 MacBookPro. Last week I tested Google Gemma 4 E4B on the same 16GB M1 MacBookPro, and the week before that it was Ternary Bonsai 2, the 2-bit squeeze of Qwen 3.8 27B.

If you want to jump to the video, go ahead:

If you prefer reading, let’s go.

Nothing changed in the methodology, I’m still using the same questions and coding challenges, and the screen recordings are not sped up or edited in any way, so you get exactly what can you realistically run on an entry-level machine, at real speed – with the wifi off if you want it off.

Today’s model is Xiaomi’s MiMo V2.6, a 9 billion parameter distill of Qwen. Xiaomi took Qwen 9B, post-trained it on their own traces, and put it out as an agentic checkpoint. On paper it is better at long tasks and at coding that needs tools.

Fair warning that I changed one thing in the setup. In the Bonsai video I used a terminal. This time I ran everything inside MLX Core, a native Mac app that sits on top of MLX, link at the end of this blog post. It looks a little like ChatGPT or Codex Desktop, which makes the recording easier to follow. It also adds a bit of memory overhead, so the tokens come out a little slower than a raw CLI run.

The first prompt was: “what is a dolphin?”, and the answer came back in about 30 seconds, at roughly 12 tokens per second. The answer was fuller than Gemma’s and Bonsai’s. More species, more structure, it felt like a model that likes to finish the thought.

The second prompt is the small utility I keep using as a test: “write a Python script that counts the words in every file in a given folder”. MiMo did something the others did not: before it even wrote a line of code, it asked me to enable tools. It already knew it would save something on the disk. Then it thought for about four minutes and produced word_count.py. The script runs ok, but the printed output has a little garbage in it — a “mojibake in files” line that came out of nowhere — but the job got done.

I would put MiMo 2 between Gemma 4B and Ternary Bonsai 2. Maybe on the same shelf as Bonsai. A little slower, maybe. Good enough for summarization and for the kind of quick utility you can throw at a 9B model and expect to walk away with a file. I would not start a long agentic coding session on this machine with this checkpoint.

And then the part that still matters more than the tokens-per-second number: it is free, it does not bill you per word, and it keeps working when the laptop is offline.

If you have another model that you wanna see it tested on this same 16GB M1, leave a comment under the video. I already have the next three from the last round of comments.

A few links, if you want to run the same thing:

Thanks for watching and spreading the word. Every share makes local, sovereign AI a little less of a lab toy.

Previous