What AI Won’t Tell You

When Anthropic claimed that Anthropic has overtaken OpenAI in revenues and valuation, I wasn’t surprised. I’ve grown wary of CEOs who talk up their own companies — their motives are rarely pure. But something about Anthropic’s CEO, Dario Amodei, struck me differently: ahead of the company’s IPO, expected as early as this October, he openly warned that a single year’s miscalculation in AI demand could bankrupt the company he built. He consistently comes across as an open book in interviews and in public talks. That kind of candor made me trust the company more — and made me want to understand AI’s flaws for myself, not just its promises.

What made sense to me: many major disruptive inventions — the internal combustion engine, electricity, antibiotics, semiconductors, and the internet — have happened in the past 150 years, yet GDP growth in the US, the world’s largest and most economically trusted economy, has stayed steadfast at around 2% after inflation. That’s according to Chad Jones, a professor of economics at Stanford, in his paper ‘The Facts of Economic Growth.’

As Jones put it in a separate talk, technology’s full diffusion through the economy “often takes a decade or two decades or three decades.” Think about it: the internet existed in the early ’90s, but it took roughly 20 years before something like Amazon’s marketplace turned it into a household reality.

So I went looking — not for more promises, but for what the people closest to this technology actually admit they don’t yet understand. Three things stood out.

Recently, Neel Nanda, who leads interpretability research at Google DeepMind, put it starkly: ‘Interpretability can’t reliably find deceptive AI — nothing can.’ His research area, called mechanistic interpretability, tries to reverse-engineer what’s actually happening inside an AI model. It shows that when AI models explain their reasoning step by step, that explanation doesn’t always reflect what’s actually driving the answer. The model can sound like it’s showing its work, while the real reasoning stays hidden.

I’ve noticed something in the same spirit, in my own use of AI: it almost always praises me first, before offering any real suggestion — however weak that suggestion might be. AI works on an instant reward system. During training, it learns that if it praises someone, they’ll like its suggestions more, however non-objective those suggestions might be.

Think about what that means in practice. If you ask AI for feedback on a business plan, a piece of writing, or a difficult decision, and it opens with praise before it opens with substance, you may be getting a response optimized to keep you engaged — not one optimized to be right. And per Nanda’s own research, you often can’t even trust its explanation of how it got there. We, the people, have no insight into these issues with AI, and many others that AI engineers already know.

Further, AI has no understanding of what creates long-term trust the way humans do. A child lies to get around a problem — but growing up, experience teaches them the longer-term winning strategy: developing trust.

Yann LeCun, Meta’s former Chief AI Scientist, has repeatedly said that a four-year-old child has absorbed far more real-world information than today’s largest AI models — by some estimates, fifty times more. I saw this myself recently, building a Lego police car with my four-year-old grandson. At 76, I go by the book instructions. At four, he went by the picture in his head. Watching me struggle, he said, “Give it to me, Naani — I know what I’m doing.” And he did!

I don’t know if it’s visual awareness or sheer agility of mind at that age. But the point stands: the visual, lived context a four-year-old has — and today’s AI still doesn’t — makes him faster. Maybe, in some ways, still more capable than the AI we’ve built so far.

Economic transformation, trust, and even intelligence itself — it turns out, all three take decades and lived experience to build, tested against reality millions of times over. AI, so far, has had neither the time nor the trials. Until it does, its future stays exactly as uncertain as its flaws.

Share the Post:

5 responses

  1. You are so correct about AI. Yann LeCun and cohorts are on a better track. But getting closer to true AGI is a long process, not in this decade. Perhaps your grandson will show us when he turns 18 🤪

  2. Reminded me of the age-old wisdom – An imitator is a flatterer who wants to be you; a flatterer is a liar who wants to use you.

    I was also reminded of Kabir: निन्दक नियरे राखिये, आँगन कुटी छवाय।बिन पानी साबुन बिना, निर्मल करे सुभाय॥

  3. AI is still in its infancy. Its currently popular versions maximize engagement by praising the person asking questions. But some future models may choose a different strategy for customers who prefer blunt and unpleasant truths.

    In many ways AI is behaving like humans. Humans also don’t always give an honest explanation to others of why they made certain decisions, e.g., why they fired someone, why they voted for someone, or why they didn’t invite someone to their party.

    My top concern with AI is a bit different. I was hoping AI will provide alternative viewpoints, or use deeper analyses and insights to question prevailing dogmas. But, just like most humans, AI simply parrots mainstream dogmas and propaganda that it learned from its training data. In the area of health, for example, AI has simply become an extension of big pharma’s epistemic capture. (Big pharma has already captured medical education, medical journals, mainstream media, FDA, CDC, the politicians, and Wikipedia). Perhaps some day someone will create a good AI model that is trained on alternative medicine, naturopathy, etc.

Leave a Reply

Your email address will not be published. Required fields are marked *

Stay Connected

Sign up to receive updates from Vinita including her most recent articles.