really good ai model
ShareByLight

I just launched ShareByLight—transfer files with light only. No Wi-Fi, no Bluetooth, no cables. One device shows a stream, the other reads it with the camera. Free on iPhone, iPad, and Mac.

The pace at which neolabs are launching AI models is getting ridiculous. Every week is “the best week in AI”. Every few days a new “really good AI model” is launched. Every few days a new iteration on that pops up and every couple of weeks an entirely new class of models is available.

It’s becoming hard to tell what model is actually the best one. Or, put it more bluntly: what good is this AI for.

If you want to see this rant as a video, there you go:

If you prefer reading, keep going.

Honestly, I don’t believe the hype. I don’t think everything is being replaced at the speed described. I think we are getting hit by a continuous flux of information from three sources, and those three sources account for most of the new-stuff stress.

  1. paid shills
  2. neolabs PR departments
  3. red-pilled AI people

First, the paid shills. I’ve seen this many times in crypto. These are just influencers for hire, getting paid to pass the neolabs messages over.

Second, the neolabs themselves. They keep pushing a narrative that gives the impression something important is happening, or it’s just about to happen. Like yet another OpenAI model “escaping confinement” and hacking some random site on the Internet.

Third, and this one is the most impactful, just normal people who got red-pilled by AI. They are almost always genuine. They truly believe “their” model is the best. Let’s not forget that we are, essentially, just apes with clothes, so tribalism is still running high in our blood.

The thing is Opus 5.5 and Grok 4.7 are almost identical. If there is a difference, it is half a percent on whatever general scale we pretend we have.

How We Make an AI Model a “Really Good AI Model” – And Even Pay for This in the Process

We came a long way from early ChatGPT 3.0, when there was a chat box and the thing on the other side was literally some software giving plausible answers – not a call center in India. That was a real breakthrough. It bended our usage patterns so we started to get more of it. We started to ignore how often it hallucinated, and we gave extra weight to the the times when it didn’t. We walked the path that looked promising, evolving, amplifying.

And that’s how the neolabs closed the loop. They looked where we’re heading, they took our feedback and trained toward our expectations. They literally did product discovery with us as the research panel, while we paid for the privilege.

You see, different clusters of users prompt their models for different problems, so neolabs track these patterns and find verticals inside a field that is literally endless.

For instance, Grok 4.6 is very good at coding. Most of the problems that ecosystem solves are related to coding. The user base around SpaceXAI is programmers.

Gemini is very good at summarization. They did AI snippets in search for so long, that they became really good at it. The model got optimized for a tight, targeted answer that can sit in the first position, above the actual results.

Claude noticed that a lot of their customers were designers and optimized for that. Then Anthropic launched Opus 5.5 and someone made a seed movie showing what a beautiful marketing movie that model can do. If that guy was paid or genuine, doesn’t matter. What matters is that thousands of red-pilled AI people followed. Suddenly the internet is full of similarly looking videos you can glance at and immediately recognize them as Opus 5.5.

And from that one use case, people infer Opus 5.5 is way better than Grok 4.7.

It isn’t. That’s how the models are trained and exposed.

The Real Moat Is the Question

If you know what to ask, you’re suddenly in a place where pretty much every model does the job. Large language models are just statistical approximation machines. Some are trained specifically for marketing movies, some for code, some for snippets. But at the end of the day they are more or less equal depending on the question.

For example, I have been writing code for 35 years. I know what to ask when it comes to coding. I do not know what to ask about marketing or movie-making, because I hardly honed my skills in those areas. But the coding hard skills, trained over decades, allow me keep my AI spending low, because I already know that, deep down, they are all the same machine.

If you know what to ask, AI is pretty much the same. Don’t get fooled into believing Opus 5.5 is in any way better than ChatGPT Astra, or Grok 4.7, or DeepSeek, or GLM, or Qwen.

Previous