Taking Their Word For It

AI, Covid, and what happens when we start checking the experts

For as long as any of us have been alive, you trusted the expert. You went to the doctor, or the solicitor, and they told you what was what, and you took their word for it. You didn’t really have a choice. You didn’t know enough to argue, you’d no way of checking, and questioning the professional wasn’t the done thing anyway. So you nodded, and got on with it. That deference held a surprising amount together: a kind of trust we’ve handed to institutions for decades without thinking about it much.

That’s starting to come apart, and not only because of AI. Faith in the experts has been wobbling for years - Covid did a real number on it, with the whole country watching them contradict each other and change their advice in real time. Although that cut both ways. Plenty of the people who decided to ignore the experts altogether didn’t exactly cover themselves in glory. Sometimes the expert really is right, and the confident bloke online really isn’t.

AI does something different, though. It lets you actually check. You come out of the appointment, go home, and ask it what it reckons about what you were just told. Loads of people are already doing exactly this. It feels like a small thing, but it isn’t: you’re second-guessing the professional in a way that simply wasn’t available a few years ago, for free, on your phone, with nobody to feel awkward in front of. And some of the time it’s genuinely useful. Experts get things wrong, the symptom waved away or the advice that was a bit off, and now you can at least ask the question you’d never have known to ask.

It isn’t all good, though. There’s something a little corrosive about a world where nobody quite trusts the professional any more, where the doctor who’s right ninety-nine times gets fact-checked on the hundredth. And the thing we’re checking with isn’t exactly trustworthy itself; it’ll give you something completely wrong in the same confident voice it uses for everything.

The AI itself is built partly out of doctors, and not average ones. The companies behind these things pay top people in each field, leading doctors and lawyers and the rest, to sit and write the answers that train the models. So when you check your GP against it, you’re in a way reaching past them to the best in the field. That’s not nothing. But you get it all as one flat, confident summary, with the bit that tells you how much to trust any given part of it taken out.

I don’t think there’s a clean answer here. It’s a good thing and a bad thing at once. We spent a very long time trusting the people in the room because we had no real alternative. Now we have one, of sorts. I just can’t tell yet whether it leaves us better off, or simply more alone with decisions we were never that well equipped to make in the first place.

Open article

Buying the Badge

Ferrari's Luce, EVs and the one thing a Tesla can't sell

Good morning,

A contrarian take to go with the Jag’s — this time on Ferrari’s much-maligned electric Luce.

Well. I commented on the Jag with a contrarian take, so it feels only fitting that I do the same for Ferrari’s effort.

The Luce’s reception has been brutal. The share price fell close to 8%, Luca di Montezemolo, who ran the company for twenty-three years, said they should take the prancing horse off it, and even Matteo Salvini, the Deputy Prime Minister and a reliable booster of Italian industry, would not defend it. The design has been compared to a Honda Accord, a Nissan Leaf, and a Tesla Model 3 that someone sat on. The general view is that Enzo would be turning in his grave.

He probably would be. But the styling is probably the least interesting thing about this car. What the Luce is designed to do is break a rule that has held fast for a decade, and the clearest way to see it is to look at the company going the other way.

Ben Thompson made the case this week that Tesla proves something true of all modern technology: the rich and the rest end up with the same thing. The richest people he knows, he wrote, drive Teslas not for the finishes, which are cheap, but for the self-driving software, which nothing else can match. But the car a billionaire buys to drive for them is the same one half the mums in London use for school drop off.

The principle runs through everything with a chip in it. Tim Cook does not own a secret gold iPhone; he owns the one you own. There is no executive edition of the AirPods, no billionaire’s iPad. Whatever the best phone or pair of headphones happens to be, the wealthiest person alive owns the same one you can. Technology is the one luxury money cannot upgrade.

Cars were never that. A car is the most conspicuous thing most people will ever own, and the gap between the cheapest and the most expensive has always been vast and deliberate. But an electric car is barely a car in the old sense. Strip the body and it is a battery, some motors and a great deal of software, and software levels everyone. These aren’t cars; they are computers with wheels.

So for a decade the electric car turned as egalitarian as the phone. There were grander electric cars – a Rolls-Royce Spectre, a Maybach EQS – but the Rolls is built to waft and the Maybach to be chauffeured; neither is the car you buy as an ultra-premium daily driver. The best electric car you could drive hard was a Tesla Plaid, and a Tesla Plaid is what a Silicon Valley software engineer buys. If you were seriously rich and wanted to drive something electric, your shortlist was the same as someone earning a tenth of your salary: the Tesla, a Taycan, a Lucid or an electric BMW.

This is a real problem for a particular kind of buyer – the kind who cares a great deal about the statement his car makes. He might well even have a V12 Ferrari in the garage for the weekend, but he wants something electric for the week, and right now the electric car in his driveway is the same one his head of engineering parks next to his at work. He does not want that. He has never wanted that. For everyone else the levelling of the car is a curiosity; for him it is a daily irritation he has the money to fix and, until now, no way to.

The Luce is Ferrari’s fix: the first fast electric car that also makes a statement. For the first time, money buys you a better electric car to drive – and that, not the styling, is the entire point of it.

And – this is the part the styling row has drowned out – it is apparently a genuinely extraordinary car to drive. There is a motor on each wheel and around 1,050 horsepower between them; 62 mph arrives in 2.5 seconds, 200 km/h in under seven. But the raw figures are the least of it. With no conventional gearbox, Ferrari hands the driver two paddles that do something no combustion car ever could: pull the right and the four motors release their full torque for a corner exit, pull the left and they trade it for regenerative braking – the control unit re-aiming all of it two hundred times a second to keep well over two tonnes flat and gripping. It is torque vectoring of a precision a mechanical drivetrain cannot come close to.

And where almost every other maker now pipes a synthetic engine note through the speakers to paper over the silence, Ferrari does the opposite: it amplifies the Luce’s own sound, the real frequencies of its motors and electronics, and leaves it there. It is a small thing that says a large one. This is not a company faking the old Ferrari; it is a company deciding what the next one sounds like.

The objection to the Luce comes in two parts: that it is ugly, and that it is not “a Ferrari”. It is both. But the two have different causes, and only one of them is Ferrari’s fault. Start with the part that isn’t. A combustion Ferrari is dramatic because it is allowed to be; an electric one is not. This is a piece of technology, and the technology dictates the shape. At around 2,300 kilos the Luce is the heaviest Ferrari ever made, heavier even than the Purosangue, and most of that weight is a slab of cells laid flat under the floor. That mass fixes the proportions before a designer draws a line: the long wheelbase, the high body, the short overhangs. Then there is the air. Range is the only number that matters in an electric car, and at speed the enemy of range is drag. The Luce has the lowest drag coefficient of any Ferrari ever made, 0.254, against roughly 0.34 for the F40. Everything that makes a combustion Ferrari beautiful – the low nose, the deep intakes, the splitters, the wing – exists to manage air at speed, and every one of them creates drag. None of it survives the switch to electric, and that was set before Jony and Marc first put pencil to paper.

Ferrari’s chief designer, Flavio Manzoni, said it plainly: on a combustion car you design for downforce, on an electric one you design for drag, and the body “must be almost continuous, almost flat. Even the wheels must be flat.” That is the constraint.

It is why fast electric cars have all drifted toward the same general silhouette – the Tesla, the Lucid, the Taycan, the Mercedes EQS, each with the same big, long fastback, the same flush wheels, the same smooth, intakeless face. They did not copy one another. They were designed by the same maths. Until battery chemistry improves1 enough that range no longer dictates the shape, every fast electric car will look broadly like this one.

This is also the answer to the obvious objection, which is that plenty of electric cars look good. Some do. But the most dramatic ones – the Chinese electric supercars, say – manage it by not caring about range, a freedom a car meant to be lived with does not have. The handsome ones that are meant to be lived with – basically only the Taycan – are handsome within the same smooth envelope, not in defiance of it. Drama and range pull in opposite directions, and Ferrari built the Luce to be driven on the road, not the racetrack.

The Luce cannot win on looks, then, and it does not need to. It is not competing with the F80, the 12Cilindri, or any car with cylinders – Ferrari has no plans to stop building those. It is competing with the Tesla and the Taycan, where it holds an advantage no rival can engineer. A Porsche is a driver’s car; a Ferrari is a statement, and at this level the statement is what the money buys.

And then the same logic, multiplied, points east. The buyer I have described exists in London and California, but there are far more like him in Shanghai, where Ferrari has never properly cracked the market and no one feels much nostalgia for a V12 they were never really sold. Nor is there much patience there for quiet luxury: the Luce wears no fewer than eight prancing horses, with a light-up key as a ninth – the kind of thing that “accidentally” falls from a pocket onto a restaurant table. In the land of the logo, that is a feature, not an embarrassment. The tax seals the deal. China punishes big engines – a displacement levy reaching 40%, before import duty and VAT – potentially doubling the cost of the car – so a combustion Ferrari costs far more there than at home, while an electric one skips the engine tax entirely.

The other charge – that it is simply ugly – is the fair one. Ive and Newson’s rounded minimalism suits a phone, a thing you are meant to forget you are holding; but a $500,000 car is meant to be looked at. The colours are loud, but the surfaces beneath them are forgettable, and forgettable is the one thing an object bought to be seen cannot be. The envelope was forced on Ferrari; the dullness was not. They had room to do more within it, and didn’t.

Tesla killed off the Model S and X this year and handed the factory to its robots; its richest customers now have the same cars as everyone else. Ferrari has gone the other way, and built the rich a car the rest of us cannot have. The Luce is not really selling a car. It is selling the one thing a Tesla cannot: the badge.


1. The chemistry that will eventually change that – solid-state cells, silicon anodes – is a good part of what we focus on over at CanaryIQ.

Open article

Transparent Aluminium

AI is mapping every material we could have. Making them is another matter.

There’s a scene in Star Trek IV - the one where they go back to the eighties to fetch a couple of whales - where Scotty needs a huge tank built and hasn’t got the materials for it. So he sits down at a clunky old Mac, tries talking to it, gives up, picks up the mouse like a microphone, and in the end just types out the formula for “transparent aluminium” in exchange for the plexiglass he’s after. The manufacturer stares at the screen, baffled, while Scotty cheerfully hands over a material that won’t be invented for a few hundred years.

That scene stuck with me as a kid, more than the phasers or the warp drive ever did. Something about the idea that the really futuristic thing wasn’t a machine at all, but a material - a substance we just didn’t have the recipe for yet.

I keep thinking about it lately, because we’re getting close to the first half of that. Not the chatbots, but a different kind of AI - the sort that models physics and chemistry rather than language. It’s started working through the space of materials that could exist: which combinations of elements would actually hold together, which crystals are stable, which steels and glasses are perfectly possible but nobody has ever happened to mix. The numbers are staggering. People are talking about hundreds of thousands of stable materials we’d never catalogued, found not in a lab but on a computer, by working out what nature will and won’t allow.

So the recipe part is, slowly, arriving. We’re going to know what’s possible.

The catch is that Scotty didn’t just name the stuff. He had to hand over how to make it, and even with the formula in front of them, that eighties manufacturer would have spent years and a fortune actually producing it. That’s the part the film skips over, and it turns out to be the part that matters. Knowing a material is possible is nothing like being able to make one. Predicting that something is stable is the easy bit next to working out the temperatures, the pressures, the order of the steps, the fiddly craft of getting atoms to line up like that out in the real world.

And, as it happens, transparent aluminium is real. There’s a ceramic called ALON that gets sold, fairly straight-facedly, as “transparent aluminium armour”. The reason you’re not looking out of an aluminium window right now isn’t that nobody knows it’s possible. It’s that making the stuff is slow, hard and expensive.

Which is roughly where I think we’re going: a growing list of miracle materials we know are possible and mostly can’t make yet. A library of the future, running miles ahead of the factory. The hard part stops being what could exist and becomes the duller question of how on earth (or perhaps not) you actually build it. Scotty got the formula across in about a minute. The rest of it was always going to be the real work.

Open article

26x

Mind. Blown.

I’ve started telling people that I am working at 26x speed. It sounds like hyperbole, which is sort of deliberate. But it’s not. Since the beginning of the year, I have done six months of work – of coding – every week.

Working with AI agents, and especially Claude Code, has changed my entire coding life.

I’ve always been driven to build, and describe myself as a builder on my personal page, but I’ve often become frustrated with how long it takes to do the things required to build something great. Even a simple CRUD app takes a decent amount of footwork. Add an item. Delete an item. Edit an item. Add validation. Build relationships between items. That’s hours of work.

I would get bored implementing all of the basic states. It would run my battery down, even as I pushed through. The incremental dopamine hits of a function working as I hoped were sometimes enough – but there came a point where I just couldn’t do any more for the day. And I consider myself to be pretty persistent.

5x–10x

Then came the first “good” versions of LLM code writing. Cursor. It would autocomplete, correctly, what I wanted to do more often than not. I started chatting and asking it to build out features, iterating back and forth until we got to where we wanted to be.

By the middle of last year I had become the coder I wanted to be. The tooling had made it possible for me to complete around a week of 2024 work in a day – albeit with the caveat that most of my time was spent checking the work, with occasional bouts of huge frustration as the LLM did something completely inane.

Codex was a game changer. Moving to Codex from Cursor meant that I could finally level up to being a 10x coder. Two weeks of work in a day. We need a micro app to solve a very specific internal problem? No worries, just give me a day.

My output was high, but I still sank a lot of my time into making sure that the code was right and really tweaking designs. It was great, but it wasn’t great.

And then came Claude

I had used Claude Code a little previously but had stuck to the OpenAI ecosystem – the tools were pretty comparable. Until an update in December, when everything changed. I saw murmurings on X and decided to play a little over the Christmas break. It was an entire reality shift.

Spec out what you want built, and it would build it. End to end. No laziness. Correct integrations. It would actually go and get the documentation I asked it to. It removed all of the effort involved in building an app.

Ideas that have sat on the shelf for years are built and live. In the course of a month I have done more than I could have done in well over a year. It has blown my mind. My problems are no longer code related. They are all actual business problems.

This year I will produce more code – more features – than I could have been expected to in my entire career just five years ago. Which is just… wild.

Open article

The Models Don't Matter Anymore

Frontier releases keep coming, but for businesses, focusing on flexible infrastructure pays bigger dividends than chasing the latest and greatest.

Chop Chop!

  • Model-Agnostic Infrastructure Build systems that can swap between LLMs seamlessly, ensuring quick upgrades and minimal disruption when new models arrive.
  • Focus on Data Quality Clean, well-labeled, and accurately governed data can make even mid-range models shine, while poor data can sink any AI effort.
  • Fine-Tuning Is Overhyped Constantly emerging frontier models often outpace costly fine-tuning.
  • Prompt Engineering Matters Well-crafted prompts can embed domain knowledge and outperform heavily customized models—often with minimal overhead.
  • Prepare for Rapid Change Software and strategies that take 18 months to deploy risk obsolescence. Design for quick iteration and disposable components.

In Full

Another day, another frontier-level release—welcome to the world, Claude 3.7 Sonnet. And yet, the churn of model announcements increasingly feels like background noise. For the past two years, conversations about AI have fixated on which model is best: who has the most parameters, who trained on the biggest cluster, whose benchmarks are highest… But do these improvements really matter anymore?

Ordinarily, when an article poses a rhetorical question, the safe bet is: “No.” But that’s not actually the case here—new models do matter. Yet for most businesses and consumers, the impact today is modest at best.

The constant stream of model releases are the result of the “great race”. OpenAI, Google, Anthropic, xAI, Mistral, DeepSeek, et al. are all chasing what they hope will be the model—the one capable of AGI (Artificial General Intelligence). Definitions vary. OpenAI’s charter describes AGI as systems that “outperform humans at most economically valuable work.” Sam Altman recently tweaked this definition to something like “a system that can tackle complex problems at human level, across many fields.”

My own definition is more specific:

The first true AGI is the first model that can, with limited human input, build a better version of itself.

In fundamental research then, the drive for better models is far from meaningless. Performance gains, more elegant architectures, and greater efficiency keep nudging us closer to Altman’s definition of AGI—if not OpenAI’s or my own. Specialized models also matter, and matter a lot: Google’s “co-scientist” model has shown real promise (paper), replicating in days hypotheses that scientists themselves honed over a decade of research (blog).

But for most businesses—and for most of us—this doesn’t (and shouldn’t) carry that much weight.


Don’t Let Perfect Be the Enemy of Good

When LLMs first arrived, the experience was mind-bending. The GPT-3.5 generation could generate coherent prose so convincing that a Google engineer believed their system had achieved sentience. It was natural to see those emergent abilities as a quantum leap in machine intelligence, because, well, they were.

The jump from GPT-3 to GPT-4, Claude 3.5 Sonnet, Google’s Gemini, and the like felt like another quantum leap. With the right coaxing, they became really good. Now, in the newest wave—OpenAI’s “o-series,” DeepSeek’s R1, and Grok 3—we again see models that are …really good. The leaps are getting harder to spot, not because they aren’t happening but because for most tasks and most people, existing LLMs already do the job.

For standard day-to-day tasks that require college-level writing or problem-solving, today’s models meet the bar. The difference between, say, Claude 3.7 Sonnet and what might be released next month is real, but often subtle. The game-changing improvements have already happened.

My argument is that the current crop of models are the ones businesses should build around—regardless of what’s coming next, and even if they don’t always beat human performance right now. By the time your project goes live, the model you spec’d on Day 1 will be replaced by a faster, cheaper, and more capable successor anyway.


Preparing the Groundwork

The key is to be model-agnostic. Treat the model like a new employee who will be “upskilled” every few months. Don’t design systems that rely on the quirks of a single LLM; assume the best option will shift frequently.

I’ve spent the past couple of months thinking about general rules for working with AI:

1. Build to Be Model-Agnostic

It’s safe to assume that models will keep improving and that different models excel at different sub-tasks. Your system should let you swap between them easily, upgrade seamlessly, and test performance quickly. Create well-structured pipelines so you can deploy cheaper models for high-volume tasks, then tap a more powerful model only when complexity demands it. It’s the same principle we use at CanaryIQ: we allocate tasks to whichever model handles them best. This flexibility is crucial given how fast the AI landscape evolves and how new capabilities appear “overnight”.

2. Focus on Your Data Integrations

Data remains the ultimate determinant of AI success. Even the most advanced model will falter if your data is inaccurate, incomplete, or poorly labeled. Conversely, an average LLM can shine with well-structured and reliable data. Many organizations discover that data engineering is the real bottleneck. Cleaning, labeling, and governing your data often yields bigger returns than anything else–and this can be done using current gen LLMs.

3. Fine-Tuning Models Is (Probably) Not Worth It

Once upon a time, fine-tuning was the standard for customizing a model to a specific domain. But in an era where a new, frontier-level release appears every few months, that strategy can quickly become obsolete. Why spend loads of engineering effort and dollars on a fine-tuned model that might be overtaken by a more advanced general model tomorrow?

As context windows expand, you can simply feed more relevant data into the prompt itself, sidestepping the need for complex retraining. This approach makes it easier to pivot as soon as a better model arrives.

4. Get Good at Prompt Engineering

Effective prompt engineering can embed domain expertise without entrenching your workflows in a single model. In many cases, LLMs themselves can help generate optimized prompts for specific tasks. You can automatically iterate through variations, test performance, and integrate the best prompts based on internal benchmarks. By treating prompt engineering as an ongoing process, you keep your system nimble and resilient.

5. Orchestration, Memory, and Retrieval-Augmented Generation

Real-world tasks often go beyond a single question-and-answer step. You may need to retrieve documents, compile them into a prompt, verify outputs, or query secondary data sources. This multi-step chain requires orchestration frameworks that let you slot in different LLMs as needed. If your logic isn’t bound to the quirks of one model, it’s easy to upgrade whenever a better contender shows up.

6. Agentic Workflows and Tool Use

Automated “AI agents” that chain tasks, call APIs, and pursue complex goals are the biggest news at the moment. But they’re only as useful as the systems behind them—reliable data access, robust orchestration, and well-defined processes. The choice of LLM matters far less if the rest of your infrastructure is brittle.

7. Models as Commodities

LLMs are commoditizing. The leading-edge model might shine for a particular challenge, but only until the next iteration arrives. You can’t simply chase “the best” model indefinitely. Instead, build flexible platforms that let you test new releases, confirm improvements, and switch as soon as it makes sense.

8. Disposable Software

The AI market moves at ludicrous speed. There’s little point in an 18-month build cycle if everything is outdated the day you ship. This doesn’t mean writing sloppy code, but it does mean building in a way that expects short life cycles. Your architecture should let you drop in a new LLM, run tests, confirm upgrades, and go live with minimal reengineering.


Looking Ahead

Models will continue to matter—and they’ll keep hitting the front-page of hacker news. At some point, we may even cross the threshold of AGI by my own definition (a system that can iteratively improve itself) and all hell may break loose. But for most organizations, whether that milestone arrives next year or the next decade is noise.

Focus on what you can control: a flexible infrastructure that pivots whenever a better model appears. That’s the essence of future-proofing your AI strategy. Don’t obsess over which LLM is hot this month—just be ready for whatever’s next. The models keep changing, but the right structures will outlast every frontier release.

*This article came from a talk I’ve given on this subject - the last version is here

Why It Matters

Because the AI market moves faster than any single model’s development cycle. Building around a specific LLM risks missed opportunities as better options emerge. By embracing a model-agnostic strategy, you future-proof your AI efforts, reduce implementation headaches, and keep your focus on what truly drives results: data quality, flexible infrastructure, and the ability to pivot as soon as new capabilities arise. In short, the real competitive advantage doesn’t come from chasing the best model—it comes from designing systems that can absorb whatever comes next.

Open article

Earlier

Feel like things are changing faster than ever?

You're right.

Get each new piece by email, the moment it's published. Delivered by Substack, free.

Already a subscriber? Read on Substack.