Modern Australian
The Times

AI is failing ‘Humanity’s Last Exam’. So what does that mean for machine intelligence?

  • Written by Kai Riemer, Professor of Information Technology and Organisation, University of Sydney
AI is failing ‘Humanity’s Last Exam’. So what does that mean for machine intelligence?

How do you translate ancient Palmyrene script from a Roman tombstone? How many paired tendons are supported by a specific sesamoid bone in a hummingbird? Can you identify closed syllables in Biblical Hebrew based on the latest scholarship on Tiberian pronunciation traditions?

These are some of the questions in “Humanity’s Last Exam”, a new benchmark introduced in a study published this week in Nature. The collection of 2,500 questions is specifically designed to probe the outer limits of what today’s artificial intelligence (AI) systems cannot do.

The benchmark represents a global collaboration of nearly 1,000 international experts across a range of academic fields. These academics and researchers contributed questions at the frontier of human knowledge. The problems required graduate-level expertise in mathematics, physics, chemistry, biology, computer science and the humanities. Importantly, every question was tested against leading AI models before inclusion. If an AI could not answer it correctly at the time the test was designed, the question was rejected.

This process explains why the initial results looked so different from other benchmarks. While AI chatbots score above 90% on popular tests, when Humanity’s Last Exam was first released in early 2025, leading models struggled badly. GPT-4o managed just 2.7% accuracy. Claude 3.5 Sonnet scored 4.1%. Even OpenAI’s most powerful model, o1, achieved only 8%.

The low scores were the point. The benchmark was constructed to measure what remained beyond AI’s grasp. And while some commentators have suggested that benchmarks like Humanity’s Last Exam chart a path toward artificial general intelligence, or even superintelligence – that is, AI systems capable of performing any task at human or superhuman levels – we believe this is wrong for three reasons.

Benchmarks measure task performance, not intelligence

When a student scores well on the bar exam, we can reasonably predict they’ll make a competent lawyer. That’s because the test was designed to assess whether humans have acquired the knowledge and reasoning skills needed for legal practice – and for humans, that works. The understanding required to pass genuinely transfers to the job.

But AI systems are not humans preparing for careers.

When a large language model scores well on the bar exam, it tells us the model can produce correct-looking answers to legal questions. It doesn’t tell us the model understands law, can counsel a nervous client, or exercise professional judgment in ambiguous situations.

The test measures something real for humans; for AI it measures only performance on the test itself.

Using human ability tests to benchmark AI is common practice, but it’s fundamentally misleading. Assuming a high test score means the machine has become more human-like is a category error, much like concluding that a calculator “understands” mathematics because it can solve equations faster than any person.

Human and machine intelligence are fundamentally different

Humans learn continuously from experience. We have intentions, needs and goals. We live lives, inhabit bodies and experience the world directly. Our intelligence evolved to serve our survival as organisms and our success as social creatures.

But AI systems are very different.

Large language models derive their capabilities from patterns in text during training. But they don’t really learn.

For humans, intelligence comes first and language serves as a tool for communication – intelligence is prelinguistic. But for large language models, language is the intelligence – there’s nothing underneath.

Even the creators of Humanity’s Last Exam acknowledge this limitation:

High accuracy on [Humanity’s Last Exam] would demonstrate expert-level performance on closed-ended, verifiable questions and cutting-edge scientific knowledge, but it would not alone suggest autonomous research capabilities or artificial general intelligence.

Subbarao Kambhampati, professor at Arizona State University and former president of the Association for the Advancement of Artificial Intelligence, puts it more clearly:

Humanity’s essence isn’t captured by a static test but rather by our ability to evolve and tackle previously unimaginable questions.

Developers like leaderboards

There’s another problem. AI developers use benchmarks to optimise their models for leaderboard performance. They’re essentially cramming for the exam. And unlike humans, for whom the learning for the test builds understanding, AI optimisation just means getting better at the specific test.

But it’s working.

Since Humanity’s Last Exam was published online in early 2025, scores have climbed dramatically. Gemini 3 Pro Preview now tops the leaderboard at 38.3% accuracy, followed by GPT-5 at 25.3% and Grok 4 at 24.5%.

Does this improvement mean these models are approaching human intelligence? No. It means they’ve gotten better at the kinds of questions the exam contains. The benchmark has become a target to optimise against.

The industry is recognising this problem.

OpenAI recently introduced a measure called GDPval specifically designed to assess real-world usefulness.

Unlike academic-style benchmarks, GDPval focuses on tasks based on actual work products such as project documents, data analyses and deliverables that exist in professional settings.

What this means for you

If you’re using AI tools in your work or considering adopting them, don’t be swayed by benchmark scores. A model that aces Humanity’s Last Exam might still struggle with the specific tasks you need done.

It’s also worth noting the exam’s questions are heavily skewed toward certain domains. Mathematics alone accounts for 41% of the benchmark, with physics, biology and computer science making up much of the rest. If your work involves writing, communication, project management or customer service, the exam tells you almost nothing about which model might serve you best.

A practical approach is to devise your own tests based on what you actually need AI to do, then evaluate newer models against criteria that matter to you. AI systems are genuinely useful – but any discussion about superintelligence remains science fiction and a distraction from the real work of making these tools relevant to people’s lives.

Authors: Kai Riemer, Professor of Information Technology and Organisation, University of Sydney

Read more https://theconversation.com/ai-is-failing-humanitys-last-exam-so-what-does-that-mean-for-machine-intelligence-274620

Modern AI SEO Agency vs Traditional SEO: What’s the Difference

Search engine optimisation has changed dramatically over the past few years. Search engines have become smarter, user behaviour has evolved, and bus...

Caravan Travel for Modern Australian Getaways: Plan a Comfortable Holiday

A family road trip is one of the best ways to explore Australia together. And, travelling by caravan gives you the freedom to take your time, stop a...

Mini Excavator and Trailer Package for Sale: What I Buy as One Deal in 2026

The first client who asked me for a mini excavator and trailer package for sale wasn’t trying to save a few hundred dollars on shipping. They we...

Make Dad a Guest in His Own Home This Father’s Day

Father’s Day can accidentally turn Dad into the unpaid event manager of his own celebration. He lights the barbecue, finds extra chairs, checks wh...

Where to Enjoy Your Off-Road Caravan on the Gold Coast

With a caravan, you can travel anywhere and everywhere without battling the rush of the peak holiday season or last-minute reservations. While the r...

How Osteopathy Supports Recovery from Sciatica and Nerve Pain

Sciatica isn't just annoying. It's genuinely painful. It sits deep in your glute and shoots straight down the back of your leg. It turns something as...

The Winter Jewellery Edit: Five Pieces You'll Wear All Season

As wardrobes shift to cosy knits, tailored coats and rich seasonal textures, jewellery becomes the finishing touch that pulls every winter outfit to...

7 Signs It's Time to Upgrade Your Piston Air Compressor

If you run a workshop, panel shop, or fabrication business anywhere around Perth, you already know what heat and dust do to equipment over a few sum...

How Long Do Bathroom Renovations Melbourne Take? Step-by-Step Process Explained

Planning a bathroom renovation is exciting, but one of the biggest questions homeowners ask is, "How long will it take?" While every project is uniq...

Why Your Skin Breaks Out: The Science of Acne Explained

Acne is the most common skin condition in the world. An estimated 85% of people experience it at some point between the ages of 12 and 24, and a gro...

10 Swimwear Trends Australian Women Are Wearing This Summer

Every Australian summer brings a fresh wave of swimwear trends, but some styles have much greater staying power than others. While fashion constantly ...

Why Regular Skills Updates Are Essential for Licensed Security Officers

A guard at a Brisbane shopping centre gets a call about a shoplifter who's turned aggressive.  They’ve done the job for six years. But their de-...

10 Benefits of Choosing Professional Tutoring Penrith Services

Every student has unique learning strengths, challenges, and academic goals. While classroom teaching provides essential knowledge and structure, so...

Sunshine Coast Baby Classes Prove Big Hit Among First-Time Mums

There's a movement gaining traction on the Sunshine Coast, providing a village of support, socialisation and relief for first-time mothers and babie...

Father's Day Gift Ideas for Men Who Are Hard to Buy For

Some dads are easy to buy for. Others do not want anything, already have everything, or give you the classic "don't worry about me" answer every yea...

Top 5 Mistakes That Wear Out Your Brakes Faster

Brakes don't need frequent replacements like oil changes do.   But a lot of the wear happens quietly, over months, because of habits most drivers...

Plantation Shutters vs Curtains: Which Is Better for Your New Home?

Moving into a new home is an exciting opportunity to personalise your space and make it your own. While many homeowners focus on furniture, flooring...

Celebration of Life vs Traditional Funeral: What's the Difference?

When saying goodbye to someone you love, there is no single way to honour their life. Every family has different traditions, beliefs, and preference...