[ Video ] How To Set Up Local AI on a 16GB MacBook
Setting up local AI on a 16GB MacBook involves using optimized binaries for the Metal architecture. Start with MLX-serve. Optionally OpenWebUI
Decision tools and frameworks for autonomous living
Setting up local AI on a 16GB MacBook involves using optimized binaries for the Metal architecture. Start with MLX-serve. Optionally OpenWebUI
At the end of the day, all AI models are the same: statistical approximation machines. A really good AI model is just in the question you [...]
This week I'm testing Xiaomi MiMo V2.6, Qwen 9b distill, on a 16GB M1 MacBookPro. It works good enough for text generation and light coding
This week I'm testing Google Gemma 4 E4B on a 16GB M1 MacBookPro. It works surprisingly well for text generation and even light coding.
Grok 4.7 was launched yesterday and I had a chance to test it. It has better prompt adherence, it's more thorough, but also token greedy.
A System One model, like JEV, doesn't output text, it only decides, and that's why is perfect for my Assess Decide Do framework.
This is a brief test of a heavily quantized version of Qwen 3.8 27B, fitting in 7GB on disk. Impressive, but not 98% of the original [...]
95% of my blog is coming from agents now. They didn't ask for permission, those bots, so I decided to make the AI agents pay.
Open source AI is here to stay, so let's see what open weights models you can realistically run on a medium laptop as of September 2026.
In a rather disturbing essay, Dario Amodei, Anthropic's CEO, urges to slow down AI dissemination (not training) and ban open source AI.
It's very easy to vibe code an entire SaaS these days. But the cost of building going down doesn't mean ALL the apps will disappear.
Vibe coders are not coders. And that might be their unfair advantage. For these lazy, creative people I built a vibe coding toolkit.
You can now share your Grok bots. I had early access to the feature and I published my own before they went live. I explain how [...]
Grok Bot quick review: it looks simple, but at the same time useful, just like the first iPhone. Agents talking to each other feels magical.
Use Grok Build with local and OpenRouter free models alongside Grok-4.6. Copy-paste config.toml for Gemma, Qwen, and GLM — same harness, switch with /model.
In the final episode I talk about open weight model harnesses, or how to actually talk to your local AI. From llama-cpp chat to Pi, or [...]
In this episode I talk about running an open weights model locally. Two main hardware lines: Apple Silicon and NVIDIA, both good.
Episode 2 of the mini series about how to choose an open weights model, this time how to choose an open weights model provider.
How to choose and open weights model: the difference between slow, dumb and laggy - and state of the art inference, like Claude or Grok.
AI watermarking is not baked into model weights, it works at the generation level, by sampling different tokens, paired with a private key.
I've been blogging for 20 years
ask me anything