It's been more than a year since my last email, and my newsletter is back, now called Lior Builds. I'm starting with the thing I spend most of my day on: picking the right AI model for each job.
If we've never met: I've been shipping software since 1993. I've built products used by millions of people, like BookAuthority, Competely and TailoredRead, and they've been featured in The New York Times, CNN, Forbes and TechCrunch. Today I run Lifehack Labs, the company behind them, and I build all my products with AI.
That means I spend most of my day working with AI models: writing with them, coding with them, automating my life and business with them, and putting them inside products that real customers pay for. Every choice below comes from that work. When a model is inside one of my products, it earned its place by beating the others on the real job, not on a leaderboard or benchmark. When it’s a model I use for my own work, the choice comes from months of daily use, plus independent tests I make.
These days, new models come out every week, so I'll keep this guide updated every once in a while.
The short version (TL;DR)
My daily driver for thinking, writing and design: Claude Opus 5.5
My second opinion for reviews, math and experiments: OpenAI’s GPT-6 Astra
Checking facts before I act on them: GPT-6 Astra
Big-picture strategy reviews: Claude Fable 5.1
Writing book-length prose: Claude Sonnet 5, with thinking set to low
Text people read on a web page: Claude Opus
Tables and structured data: OpenAI’s GPT-5.6 Luna
Answering one question many times: Jev, from TypeSafe AI
Small decisions at scale: Claude Haiku 4.5, now testing Haiku 5.5
Narrating a whole book: OpenAI’s gpt-4o-mini-tts, with the Echo voice
Drawing illustrations: OpenAI’s gpt-image-1-mini
Grading AI output: a judge from a different company than the writer
Writing code: Claude Opus 5.5, with GPT-6 Astra as the reviewer
Dictation: Avalon 1.5, Aqua Voice’s own transcription model
Claude is my daily driver. GPT is the genius I consult.
Here’s the simplest way I can put it. OpenAI’s top models, like GPT-6 Astra and GPT-5.6 Sol, are like a person with an IQ of 200 and no social skills. They’re brilliant at adversarial review, statistics, experiment design and math. They find the hole in a plan that everyone else missed. But they’re mediocre copywriters, weak at marketing and design, and harder to talk to. They overengineer, they nitpick, and they cause scope and feature creep.
OpenAI’s top models are like a person with an IQ of 200 and no social skills.
Claude has taste. It writes better copy, makes better design calls, and understands what I mean without three rounds of clarification. So Claude is the model I work with all day, and I bring in GPT when I want something torn apart. Before I commit to an architecture, an experiment or a strategy, I ask Astra to find what’s wrong with it. It usually always finds something.
Claude Opus 5.5 earned the daily-driver spot recently. Its predecessor, Opus 5, came out on July 24, and I couldn’t stand it. It was verbose and yet hard to understand, it invented its own jargon, it argued with my instructions, and it sometimes said a task was done when it wasn’t. Its writing was unbearable: verbose yet compressed, dense and elliptical, trying to be clever, and too full of analogies and theatrics. I wasn’t alone: developers complained about the same things publicly for weeks. Someone even built a site to show how frustrating Opus 5 was. Opus 5.5 came out on September 22 and fixed all these issues. It leads with the answer, writes in plain words, makes fewer mistakes, and doesn’t shorten my lifespan.
Checking facts and thinking big
When a decision depends on research, I don’t act on it until a second model has checked the facts. That model is GPT-6 Astra, which made up facts less often than the other leading models in Artificial Analysis’s September test.
For a monthly and a quarterly review of the whole business, I use Claude Fable 5.1. Its job is to find connections, risks and indirect effects across everything I run, and it’s the most “big picture” of the models I use. It’s fantastic at connecting the dots across different domains and thinking outside the box. Unfortunately, it also makes up facts more often, so anything I’m about to act on from a Fable review gets checked by Opus or Astra first.
I learned to pick these models by hand the expensive way. I run 23 Claude jobs on a schedule, and I’d left most of them on the “Default” model, which was Opus. Around August 14, Default switched to Fable 5, a pricier model, and every job went with it. My usage tripled overnight. It took two weeks to notice. Now every job has a model I chose.
Writing book-length prose: Claude Sonnet 5, with thinking set to low
TailoredRead writes custom nonfiction books, so writing is the product. Claude can think before it writes, and a setting controls how much. I tested all five levels.
The default level, high, was the worst on quality, cost and speed. Sometimes it thought for so long that it ran out of room and stopped a chapter mid-sentence. Low wrote books that were just as good, for about 40% less.
The bigger Opus 5 scored higher with an AI judge, at about 2.5 times the cost per book. But bigger models have lost my own blind reading before. In May, a GPT judge scored Opus 4.7 above Sonnet 4.6. Then I read 20 pairs of passages blind, without knowing which model wrote which, and picked Sonnet 11 to 8. A week later Opus 4.8 came out, the judge preferred it again, and I picked Sonnet 12 times to Opus’s 3.
When I reread the passages I’d picked, my taste was consistent: short sentences, plain modern words, concrete scenes, named people and worked-out numbers. The bigger model doesn’t always write that way. So my blind read is now part of the eval process and every model decision. The AI judge doesn’t get the final say. Sonnet 5.5 and Opus 5.5 get the same blind test next.
Text on web pages, and the tables next to it
Packfits is a site of travel packing guides, and until September every page was written by GPT-5 mini. To be honest, GPT-5 is a lousy writer with AI telltale signs, and readers may leave a page for that reason alone. I asked two AI raters, blind, whether the pages were written by a person or a machine. They were 97 to 99% sure it was a machine. I tested nine OpenAI and Anthropic models on real page prompts, and GPT-5 mini finished last. With the same rewritten instructions, GPT-5.6 Luna still read as machine-written, 91 to 95%. Claude Opus 5 came close to human travel writing.
Each Packfits destination page has 19 sections. Eight are prose people read, and the rest become tables. In a second blind test, scored by AI judges, Opus won all eight prose sections, with 88 out of 100 against 68 for the cheapest model. Even Opus 5, a model I disliked working with, wrote the best page copy.
The tables went to GPT-5.6 Luna. It got the format right on the first try every time, at a fraction of the price. One page, two models, each doing what it’s best at.
One question for every row of a spreadsheet: Jev
This summer a new kind of AI model showed up: decision models. Instead of writing text, they answer with a category, a yes or no, or a score, plus how sure they are.
I built Rowmotive based on one of these models, Jev. You paste a list, like support tickets or survey answers, ask a question, and it answers for every row. In my tests, Jev answered one question for 10,000 rows in about 17 seconds, for about 7 cents.
It isn’t perfect. On a simple test of 400 rows, Jev matched my own answers about 95% of the time, while Claude Opus matched every time, at about 25 times the cost. So Rowmotive marks the answers Jev was unsure about, so you can focus on checking those. Rowmotive also auto-suggests which questions to ask about your table, and that job goes to GPT-5.6 Luna, the cheapest of OpenAI’s GPT-5.6 models.

Small decisions that run thousands of times: Claude Haiku
Before TailoredRead writes a book, a small model decides whether the topic needs fresh research. It runs on every draft, so the cheapest model that does it well wins. Today that’s Claude Haiku 4.5. Some of these questions are yes or no questions, or pure classification, so I’ll likely replace it with a model like Jev or OpenAI’s Decisions API soon.
Anthropic released Haiku 5.5 on October 7, at about a tenth of Haiku 4.5’s price per input token. It joins the test this month.
Narrating a whole book: OpenAI’s gpt-4o-mini-tts
I wanted to turn TailoredRead’s books into audiobooks, so I had six AI voices narrate a book for me, in the same way a professional narrator would narrate an Audible book. The voices came from five models by four companies: OpenAI’s gpt-4o-mini-tts, Google’s Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, Fish Audio’s S2.1 Pro, and Alibaba’s Qwen Audio 3.0 TTS Plus. In short clips, Google’s new Gemini voices were clearly the best. I gave one of them, the Iapetus voice on Gemini 3.8 Flash-Lite, 9 out of 10. Their price looked great too, but it was a launch discount that ends in December. At full price, a book cost $6.69 on Gemini 3.8 Flash-Lite and $10.22 on Gemini 3.8 Flash, against about $5 on OpenAI’s gpt-4o-mini-tts.
Then I listened to the whole book. About every two minutes, the narrator’s pitch and pace changed slightly. I had AI models investigate, and they found that Google’s API quietly splits any request longer than 2,000 bytes of text into separate recordings, even though Google’s launch promised hours of consistent audio. I could hear every join, and the book sounded like a bug, or like parts stitched together.
Gemini 3.8 Flash-Lite, Iapetus voice: the voice changes at 0:15, where two pieces of audio meet.
OpenAI gpt-4o-mini-tts, Echo voice, the same text: the join at 0:15 is significantly less noticeable.
OpenAI’s gpt-4o-mini-tts with the Echo voice won. It’s less expressive, but it sounds like the same person from start to finish.
Drawing illustrations: OpenAI’s gpt-image-1-mini
Every TailoredRead book gets an illustrated cover, drawn by AI while the customer watches. In May, OpenAI shut down DALL-E 3, the model that drew them, so I tested the replacements side by side on 19 real books. The newer gpt-image-2 and the cheaper gpt-image-1-mini drew equally good illustrations. But mini took about 10 seconds, and gpt-image-2 about 25. With the customer watching, speed mattered more than cost, so mini won. OpenAI retires it on December 1, so I’ll be making this choice again soon.
Grading AI output: a judge from another company
A model shouldn’t grade its own family’s work. So every AI judge comes from a different company than the writer, and GPT models grade what Claude writes. When models from both companies compete, two judges, one from each company, score everything. I read their scores side by side with my own blind reading. When the judges and I disagree, I look at why.
When a report has to be right, I go further. A model from a different company checks every sentence against the source it cites, and plain code does all the counting and math.
Writing code: Opus 5.5, with Astra as the reviewer
Claude Opus 5.5, inside Claude Code, writes most of my code. Since June, an OpenAI GPT model, running in OpenAI’s Codex, reviews every feature plan before any code is written, and it runs code review on the vast majority of code. It’s very good at finding the bug Claude missed. That’s the genius-reviewer role again.
Getting the two models to work together took a few tries. In April, while the setup kept failing, I asked for a review from the OpenAI model and got one. Something felt off, so I asked Claude who actually wrote it. Claude had written the review itself, after the OpenAI model failed.
A second reviewer also costs money. In early October, Codex re-ran a test against real AI models about 120 times while building a new Competely feature, costing me an extra $861 on my OpenAI API bill.
Dictation: Avalon 1.5, Aqua Voice’s own model
I try to build the habit of speaking most of what I type to AI, since I found it’s much faster. Not only do I speak faster than I can type, I also edit myself less when I speak. I tried a few dictation apps, and Aqua Voice was the most straightforward to use. It lets you pick the transcription model: OpenAI’s Whisper, the industry standard, or Avalon 1.5, a model Aqua built for itself. I use Avalon. The first rule in my dictation instructions tells the model never to make up words.

Talking to AI instead of typing has been an uphill battle, breaking a 35-year habit of typing, but I’m making progress. Apparently, you can teach an old dog new tricks.
Your turn
Which AI model surprised you this year, for better or worse? Hit reply and tell me. I read every one.
Get the next update
I write Lior Builds every week or two: what actually happened when I built something, the real numbers, the mistakes, and the tests that changed my mind. When a new model wins one of these jobs, subscribers hear about it first.





