Google AI


Modern Australian

AI is failing ‘Humanity’s Last Exam’. So what does that mean for machine intelligence?

  • Written by: Kai Riemer, Professor of Information Technology and Organisation, University of Sydney
AI is failing ‘Humanity’s Last Exam’. So what does that mean for machine intelligence?

How do you translate ancient Palmyrene script from a Roman tombstone? How many paired tendons are supported by a specific sesamoid bone in a hummingbird? Can you identify closed syllables in Biblical Hebrew based on the latest scholarship on Tiberian pronunciation traditions?

These are some of the questions in “Humanity’s Last Exam”, a new benchmark introduced in a study published this week in Nature. The collection of 2,500 questions is specifically designed to probe the outer limits of what today’s artificial intelligence (AI) systems cannot do.

The benchmark represents a global collaboration of nearly 1,000 international experts across a range of academic fields. These academics and researchers contributed questions at the frontier of human knowledge. The problems required graduate-level expertise in mathematics, physics, chemistry, biology, computer science and the humanities. Importantly, every question was tested against leading AI models before inclusion. If an AI could not answer it correctly at the time the test was designed, the question was rejected.

This process explains why the initial results looked so different from other benchmarks. While AI chatbots score above 90% on popular tests, when Humanity’s Last Exam was first released in early 2025, leading models struggled badly. GPT-4o managed just 2.7% accuracy. Claude 3.5 Sonnet scored 4.1%. Even OpenAI’s most powerful model, o1, achieved only 8%.

The low scores were the point. The benchmark was constructed to measure what remained beyond AI’s grasp. And while some commentators have suggested that benchmarks like Humanity’s Last Exam chart a path toward artificial general intelligence, or even superintelligence – that is, AI systems capable of performing any task at human or superhuman levels – we believe this is wrong for three reasons.

Benchmarks measure task performance, not intelligence

When a student scores well on the bar exam, we can reasonably predict they’ll make a competent lawyer. That’s because the test was designed to assess whether humans have acquired the knowledge and reasoning skills needed for legal practice – and for humans, that works. The understanding required to pass genuinely transfers to the job.

But AI systems are not humans preparing for careers.

When a large language model scores well on the bar exam, it tells us the model can produce correct-looking answers to legal questions. It doesn’t tell us the model understands law, can counsel a nervous client, or exercise professional judgment in ambiguous situations.

The test measures something real for humans; for AI it measures only performance on the test itself.

Using human ability tests to benchmark AI is common practice, but it’s fundamentally misleading. Assuming a high test score means the machine has become more human-like is a category error, much like concluding that a calculator “understands” mathematics because it can solve equations faster than any person.

Human and machine intelligence are fundamentally different

Humans learn continuously from experience. We have intentions, needs and goals. We live lives, inhabit bodies and experience the world directly. Our intelligence evolved to serve our survival as organisms and our success as social creatures.

But AI systems are very different.

Large language models derive their capabilities from patterns in text during training. But they don’t really learn.

For humans, intelligence comes first and language serves as a tool for communication – intelligence is prelinguistic. But for large language models, language is the intelligence – there’s nothing underneath.

Even the creators of Humanity’s Last Exam acknowledge this limitation:

High accuracy on [Humanity’s Last Exam] would demonstrate expert-level performance on closed-ended, verifiable questions and cutting-edge scientific knowledge, but it would not alone suggest autonomous research capabilities or artificial general intelligence.

Subbarao Kambhampati, professor at Arizona State University and former president of the Association for the Advancement of Artificial Intelligence, puts it more clearly:

Humanity’s essence isn’t captured by a static test but rather by our ability to evolve and tackle previously unimaginable questions.

Developers like leaderboards

There’s another problem. AI developers use benchmarks to optimise their models for leaderboard performance. They’re essentially cramming for the exam. And unlike humans, for whom the learning for the test builds understanding, AI optimisation just means getting better at the specific test.

But it’s working.

Since Humanity’s Last Exam was published online in early 2025, scores have climbed dramatically. Gemini 3 Pro Preview now tops the leaderboard at 38.3% accuracy, followed by GPT-5 at 25.3% and Grok 4 at 24.5%.

Does this improvement mean these models are approaching human intelligence? No. It means they’ve gotten better at the kinds of questions the exam contains. The benchmark has become a target to optimise against.

The industry is recognising this problem.

OpenAI recently introduced a measure called GDPval specifically designed to assess real-world usefulness.

Unlike academic-style benchmarks, GDPval focuses on tasks based on actual work products such as project documents, data analyses and deliverables that exist in professional settings.

What this means for you

If you’re using AI tools in your work or considering adopting them, don’t be swayed by benchmark scores. A model that aces Humanity’s Last Exam might still struggle with the specific tasks you need done.

It’s also worth noting the exam’s questions are heavily skewed toward certain domains. Mathematics alone accounts for 41% of the benchmark, with physics, biology and computer science making up much of the rest. If your work involves writing, communication, project management or customer service, the exam tells you almost nothing about which model might serve you best.

A practical approach is to devise your own tests based on what you actually need AI to do, then evaluate newer models against criteria that matter to you. AI systems are genuinely useful – but any discussion about superintelligence remains science fiction and a distraction from the real work of making these tools relevant to people’s lives.

Authors: Kai Riemer, Professor of Information Technology and Organisation, University of Sydney

Read more https://theconversation.com/ai-is-failing-humanitys-last-exam-so-what-does-that-mean-for-machine-intelligence-274620

Pool and Deck Design: How to Plan the Perfect Outdoor Living Space for Your Sydney Home

For many Australians, the backyard is where life happens. Summer barbecues, weekend swims and long evenings outdoors are all part of the lifestyle, ...

Is Solar Pool Heating Worth It? What Sydney Homeowners Should Know

There's nothing quite like a backyard pool on a hot Sydney day. But once autumn rolls in, many pools sit unused for months because the water is simp...

Planning a Luxury House Move: A Week-by-Week Timeline for Prestige Sydney Homes

Selling or buying a prestige home is a major milestone. Whether it's a waterfront residence in Birchgrove, a grand Federation home in Haberfield or ...

Downsizing or Upgrading Your Caravan? Here's How to Sell It Without the Hassle

Selling a caravan can feel like a major task, especially when you are unsure about its value, paperwork, or how to find a buyer. Whether you are dow...

The Best Overseas Adventure Holidays for Australians Who Love the Outdoors

Australia offers no shortage of incredible outdoor experiences, but sometimes the best way to satisfy your sense of adventure is to head overseas. A...

Cape Town Wine Shuttle: Winelands Tasting & Tours

Embark on an unforgettable journey through the picturesque Cape Winelands, where world-class wines and breathtaking scenery await. Our Cape Town Win...

Why Giant Rats Tail Grass Keeps Coming Back After Spraying

Giant Rats Tail Grass (GRT) is one of the most frustrating pasture weeds for farmers and lifestyle property owners. You spray an infested area, see th...

When Custom Cardboard Boxes Make Sense for Your Business

Custom cardboard boxes can be useful when a standard carton does not fit a product, packing method or presentation requirement particularly well. A ...

Virtual Livestock Fencing and GPS Tracking: Improving Visibility Across Cattle Properties

What Virtual Livestock Fencing Means for Modern Cattle Management Managing cattle across extensive properties requires more than knowing where anim...

Sydney Pawnbrokers Explained: How Hocking Your Car Actually Works

Sometimes you need cash, and you need it soon. If you own a car, you may already have a way to get it. That's what people mean when they say they've...

Moving Interstate from the Gold Coast to Brisbane (or Back)? What Removalists Wish You Knew First

Have you talked to anyone who’s done the move? They say the same thing: the drive up the M1 is the easy part. It's everything around it that catches...

Why the Spring School Holidays Are a Great Time to Visit Coffs Harbour

The spring school holidays are a good time to spend a few days on the Coffs Coast. The weather is starting to warm up, there is plenty to do outdoor...

What to Do When an Older Car Is No Longer Worth Keeping in Melbourne

Ever looked at another repair quote and wondered whether your old car is still worth the trouble? It is a common turning point for Melbourne motoris...

Your Baby's First Year: A Local Guide to Feeding, Sleep, and When to Get Extra Support

Ask ten parents in a Brisbane mothers' group how their baby is feeding or sleeping, and expect ten different answers.  Someone's baby sleeps throug...

Kitchen and Laundry Makeover Ideas That Don't Require a Full Renovation

Full kitchen renos are expensive — and most people don't actually need one.  They need the kitchen to stop looking like it's stuck in 2009, or they...

How Technology Is Reshaping the Modern Australian Commercial Kitchen

The commercial kitchen has always been shaped by technology. Refrigeration changed how ingredients could be stored, modern ventilation transformed k...

The Number on a Roller Blind Fabric That Nobody Explains

Somewhere in the fabric book, next to the colour name, there is a percentage. Three per cent. Five per cent. Ten per cent. Nobody explains it, most c...

What’s Trending in Men’s Jewellery This Father’s Day!

Finding a Father’s Day gift that feels personal, stylish and genuinely wearable is not always easy. While socks and novelty mugs have traditionally ...