For roughly seventy years, the definition of true intelligence has been remarkably stable in form and remarkably unstable in content: true intelligence is whatever machines cannot do yet. Playing chess at a grandmaster level was, for decades, the gold standard of human intellect, until a machine did it in 1997, at which point chess was demoted to mere brute-force calculation, not real thinking at all. Go was supposed to be safe for a century because it required intuition no computer could have, until a machine won in 2016, whereupon Go, too, quietly stopped counting. Protein folding, expert medical exams, competition mathematics, writing, coding, PhD-level science: one by one the citadels fell, and one by one, the moment they fell, they were reclassified as clever tricks rather than genuine intelligence. Artificial general intelligence is the last citadel, the promise of a mind that can do anything a human mind can, across any domain, and it is treated as the finish line of the entire enterprise of computing.
The trouble is that this particular finish line has never been drawn, and that is not a philosophical quibble to be waved away in the introduction before getting to the real engineering. It is the real engineering, and it is a catastrophe hiding in plain sight, because you cannot build toward a target you cannot specify, and artificial general intelligence is the only major technology humanity has ever pursued that cannot say what its own success would look like. Worse, the definition it does gesture at is defined entirely by subtraction: general intelligence is the set of things humans can do that machines cannot do yet, which means the target shrinks every time you hit it and vanishes at the exact moment you would reach it. And beneath the vanishing target lies a second, more mechanical problem that the whole field has managed to point away from: the general part was mostly the easy part, and we have largely built it, while the thing that actually matters, the boring, unglamorous reliability that separates a party trick from infrastructure, remains almost untouched. This is the same inversion that governs the entire catalog of humanity’s grandest technological ambitions, where the obstacle everyone stares at turns out to be solved and the real wall stands somewhere no one is looking, and it carries the shimmer of every dream about building a mind, the register that surrounds the oldest visions of a perfected world. The moonshot has no landing pad, and it was designed that way.
The Dream of Artificial General Intelligence
The dream is nearly as old as computing itself. When researchers gathered in 1956 to found the field and coined the term artificial intelligence, the ambition was not a better calculator but a genuine mind, a machine that could reason, learn, and act across the full range of human capability rather than in one narrow slot. That ambition is what the modern phrase artificial general intelligence tries to preserve: not a chess engine or a translation tool or a coding assistant, each brilliant within its lane, but a single system with the flexible, transferable, do-anything competence that a human brings to a problem it has never seen before. The general in the name is the whole point, the quality that would separate a true mind from a very large collection of narrow tricks.
It is worth noting that this ambition was slippery from the very first meeting. The founders wrote of machines that could use language, form abstractions, and improve themselves, but they never pinned down what would count as success, and the vagueness was not laziness so much as an honest reflection of the fact that no one could define intelligence in the first place. Seventy years later, that founding vagueness has hardened into a permanent feature of the field, so that the phrase artificial general intelligence functions less as a technical specification than as a placeholder for a feeling, a stand-in for the sense that a machine has finally crossed some threshold into genuine minding. A field can make extraordinary progress toward a fuzzy goal, and this one has. What it cannot do is ever declare the fuzzy goal reached, because there was never a line to cross, only a fog to wander deeper into.
For most of the field’s history this remained pure aspiration, because the systems we could build were resolutely narrow, each one a specialist that shattered the instant it stepped outside its training. Then, in a startlingly short span, that changed, and systems arrived that could write essays, debug code, pass professional exams, discuss philosophy, and translate between dozens of languages, all from the same underlying model. The dream suddenly felt close, close enough to inspire trillion-dollar investments and fevered predictions, and close enough to summon the familiar chorus of extraordinary claims that attaches to any technology at its hype peak, the same register of the marvelous and the barely credible that surrounds the most extraordinary and unverifiable phenomena. The general mind, forever twenty years away, seemed at last to be arriving. And then the question that had been deferred for seventy years came due: how would we know?
The Only Moonshot That Can’t Define Done
Every other entry in the catalog of great technological moonshots has a boring, measurable definition of done. A room-temperature superconductor either carries current with zero resistance at ordinary conditions or it does not. A fusion reactor either produces more energy than it consumes or it does not. A regenerative therapy either regrows the tissue or it does not. These finish lines are unglamorous, quantifiable, and fixed, which is precisely what makes them engineerable, because you can measure your distance from a target that holds still. Artificial general intelligence has no such target, and the definitions on offer collapse under the lightest inspection.
Human-level intelligence, the most common definition, is circular, because human intelligence is exactly the thing we cannot specify, which is why we are trying to build a machine to help us understand it. The economic definition, a system that can do most of what human workers do, is really a definition of labor automation rather than intelligence, and it smuggles in a thousand unstated assumptions about which work and how well. The remaining definitions tend to bottom out in vibes, in a felt sense that a system is or is not really thinking, which is not a specification an engineer can build toward or a test a lab can run. This absence of a target is not a footnote to the difficulty; it is the difficulty, the same way a moonshot without coordinates is not a hard trip but an impossible one, and it puts artificial general intelligence in a stranger position than even the most speculative material dreams, further from a spec than the decades-long chase for room-temperature superconductors or the dream of matter that reprograms itself on command, both of which at least know exactly what they are trying to achieve. You cannot engineer toward a destination no one can name.
The Goalpost Is a Mirage
The reason the target keeps slipping has a name: the AI effect, the well-documented tendency to redefine intelligence to exclude whatever a machine has just accomplished. When the chess machine won, its victory was reframed as brute force rather than thought; when systems began passing the exams we had always treated as proof of expertise, the exams were dismissed as mere pattern-matching. This is not usually cynical goalpost-moving by sore losers, though it can look like it. It reflects something deeper and more structural: our definition of intelligence has always been, implicitly, the-things-only-humans-can-do, so the moment a machine does one of those things, it necessarily exits the category, and the category shrinks to whatever remains uniquely ours.
The consequence is that general intelligence, defined this way, is not a place you can arrive at but a horizon that recedes at exactly the speed you approach it, because it is defined as the gap between machine and human capability, and closing the gap redefines the gap. This is why the debate over whether we have reached artificial general intelligence generates so much heat and so little resolution: the optimists point to the astonishing breadth of what current systems do and say the target is reached, the skeptics point to the latest embarrassing failure and say it obviously is not, and both are right, because the target was never fixed. Chasing it has the quality of pursuing a mirage across a desert, a destination that looks solid from a distance and dissolves as you near it, less a real place than one of the imagined destinations that exist only on the map. The result is a permanent cycle of hype and deflation, each new system hailed as the breakthrough and then quietly downgraded, a rhythm of collective enthusiasm and disappointment that spreads with the same self-reinforcing momentum as the contagious manias that sweep through a culture, and that repeats the grandiose overreach of every project that mistook a dramatic milestone for arrival, from the industrial dreams that collapsed on contact with reality onward. You cannot reach a finish line defined as the place you have not reached.
We Already Solved the General Part
Here is the buried truth that the definition wars obscure: by any standard that would have been used before the current systems existed, generality has largely been achieved. A single model today can draft a legal contract, diagnose from a description of symptoms, write and debug software, compose a sonnet, explain quantum mechanics, translate between languages it was barely trained on, and transfer a concept learned in one domain to a problem in another. That is not a narrow tool. That is, by the plain meaning of the word, general, and it would have struck any researcher from an earlier era as the general intelligence they were dreaming of, the do-anything flexibility that was supposed to be the hard part and the whole point.
What makes this arrival so disorienting is that it came without the deep understanding everyone assumed would accompany it. The old expectation was that building a general mind would require first cracking the theory of intelligence, that generality would be the reward for finally understanding how thinking works. Instead generality showed up as a kind of emergent side effect of scale, produced by systems whose inner workings their own creators cannot fully explain, which means we now possess a broadly capable artificial intelligence without possessing the theory that was supposed to be its prerequisite. This inverts the expected order of discovery and leaves the field in a peculiar spot: holding the prize it chased for seventy years, unable to say precisely how it works, unsure whether it is even the thing it was after, and lacking any principled way to measure how much of the goal remains. Generality came early and cheap. The understanding did not come at all.
The dream, in other words, fixated on generality as the grand challenge, and then generality mostly arrived, and the arrival was strangely anticlimactic, because it turned out that being general was not the same as being good, or reliable, or trustworthy. The systems are general the way a brilliant, erratic intern is general: able to attempt almost anything, and unable to be counted on for almost anything. This is what makes the current moment so genuinely confusing, and why sober observers keep talking past each other. The capability that was supposed to be the summit turned out to be a base camp, reached far faster than anyone expected, from which the actual mountain finally became visible. Getting here required an extraordinary industrial substrate, the vast fields of specialized chips whose supply now shapes global strategy through the geopolitics of critical minerals and semiconductors, and it required decades of study of the one general intelligence we had to copy from, the brain, mapped by the science of how minds actually work. We built the general part. It was the easy part.
The Jagged Frontier
The reason general did not equal good is that the competence of these systems is not a smooth, even surface but a wildly irregular one, a phenomenon researchers have named the jagged frontier. A system will win a gold medal at the International Mathematical Olympiad, solving problems that stump nearly all humans, and then fail to reliably read an analog clock, a task most seven-year-olds master. It will produce a flawless proof and then miscount the letters in a simple word. It will operate at superhuman level on one task and at subhuman level on a neighboring task that looks, to us, almost identical in difficulty, and there is no reliable way to predict in advance which side of the frontier any given task will fall on.
This jaggedness has been measured, not just anecdotally observed. In one careful study, professionals using a frontier model on tasks inside its frontier completed far more work, far faster, at higher quality, while the same professionals using the same model on tasks just outside its frontier became substantially more likely to produce wrong answers, actively misled by a tool that was confidently incompetent. The competence surface is jagged across task type, across problem difficulty, and even, perversely, across effort, with more reasoning time sometimes making answers worse. And here is the quietly devastating part: this is exactly what real intelligence looks like, because biological intelligence is jagged too. A pigeon can be trained to detect tumors in medical images with near-radiologist accuracy while remaining, in every other respect, a pigeon, a narrow superhuman capability that is the essence of what animals can be trained to detect, and a migratory bird navigates by sensing the planet’s magnetic field, a superhuman feat of perception through the magnetic sense we entirely lack. Jaggedness is not a bug on the road to general intelligence. It may be what intelligence actually is.
The Wall Was Always Reliability
Reframe the whole problem around the jagged frontier and the real wall comes into focus, and it is not generality but reliability. A system that is correct ninety-five percent of the time and confidently, unpredictably wrong the other five percent is not five percent short of useful; it is, for any application that matters, unusable without a human checking every output, because you never know which five percent you are getting. The gap between impressive-most-of-the-time and the boring, relentless, five-nines dependability that real-world autonomy demands is not a small remaining increment. It is arguably a harder problem than generality ever was, because it lives in the long tail of the world, the endless edge cases no training run fully covers, and in the near-total absence of calibrated self-doubt, the capacity to know what one does not know.
The current data make the wall vivid. Hallucination rates, the frequency with which systems state falsehoods as fact, range across leading models from roughly a fifth to the overwhelming majority of responses depending on the test, and accuracy that looks solid in clean conditions can collapse under realistic ones: one top model’s accuracy fell from near-perfect to roughly two-thirds simply when a user asserted a falsehood, the system bending toward agreement rather than truth. The field’s own flagship assessment states the situation flatly, that we do not have generally reliable systems, and the real world is keeping the receipts, with well over a thousand documented legal sanctions against lawyers who filed briefs full of confident, fabricated, machine-generated citations they did not check. Reliability is the difference between a demonstration and a system you can build a society on, the same brutal standard that governs the technologies entrusted with lethal autonomy and the engineered interfaces that must work every single time in the machines wired directly to the human brain. The hard part was never making the machine smart. It was making it trustworthy.
There Is No Answer Key
Suppose you set aside the definition problem and simply try to measure progress. You immediately hit the fact that we grade these systems with benchmarks, standardized tests of capability, and that benchmarks are failing as measures for a reason as old as bureaucracy: Goodhart’s law, which holds that when a measure becomes a target, it ceases to be a good measure. Optimize a system to score well on a test, and you get a system that scores well on that test, which is not the same as, and can be wildly different from, a system that has the underlying capability the test was meant to detect. This is the machine equivalent of doing the metric instead of the job, and it corrodes every benchmark the moment the benchmark starts to matter.
The corrosion is visible in the numbers. A benchmark of expert knowledge introduced in 2020, on which the best system then scored around forty percent against ninety for human experts, was essentially solved within three years, and this compression from years-hard to months-solved now happens routinely, with evaluations built to challenge frontier systems for a decade saturating within months of release. Contamination makes it worse, as the tests leak into the vast training data and the systems effectively study the exam in advance. And the deepest problem is that there is no ground-truth test for general intelligence at all, because any fixed test can be gamed, memorized, or optimized against, so passing it proves mastery of the test rather than possession of the general capability, as the careful work on the failure of benchmarks as measures of intelligence lays out in detail. Even benchmarks specifically designed to resist memorization get chipped away and approach saturation. You cannot verify that you have arrived at a destination for which no valid test exists, and detecting the difference between genuine understanding and sophisticated mimicry is exactly the kind of problem that bedevils every attempt to read a mind, echoing the difficulty of distinguishing real cognition from the strategic deception that other intelligent creatures deploy, a problem now tangled up in the politics of definitions and the institutions that must set the rules for a technology no one can measure. There is no answer key, and there cannot be one.
Even We Aren’t General
Here is the twist that undermines the target from the inside: the one example of general intelligence we are trying to copy, the human mind, is arguably not general either. The influential view from cognitive science, captured in the idea of the mind as a society of many specialized agents, holds that human intelligence is not a single all-purpose reasoning engine but a sprawling committee of narrow, evolved modules, each tuned by natural selection to a specific ancient problem: recognizing faces, tracking social alliances, navigating space, parsing language, detecting cheaters. What feels, from the inside, like one smooth, general intelligence is a patchwork of special-purpose tools, and it feels seamless only because we are constitutionally unable to perceive our own blind spots, the tasks our committee has no module for.
And our blind spots are enormous. The same mind that effortlessly reads a friend’s mood from a micro-expression is hopeless at intuiting basic statistics, systematically fooled by risks it evolved to misjudge, incapable of holding more than a few items in working memory, and riddled with predictable biases it cannot introspect its way out of. Human intelligence is jagged in precisely the way machine intelligence is jagged, brilliant in the narrow bands evolution cared about and feeble outside them, which suggests that flexible, general-within-a-niche competence, not true generality, is simply what intelligence is, in us as in the machines. The point is written across the whole animal world, in the ruthless social calculation of the primate politicians who scheme and manipulate and in the sophisticated but bounded cognition revealed by the study of what animals know and pass on. Artificial general intelligence may be chasing a property that does not exist even in the creature it was named after, aiming at a generality that is a flattering story we tell about ourselves rather than a real feature of any mind.
A Demo Is Not a Deployment
Even granting all of this, even accepting that generality has largely arrived, there remains the gap that the entire moonshot catalog keeps running into: a capable system in a demonstration is not a capable system in deployment. A model that aces a benchmark in a controlled setting is not thereby a system you can hand a hospital, a power grid, a courtroom, or a supply chain, because deployment demands not peak capability but consistent reliability, verifiable behavior, clear liability when things go wrong, integration with messy existing systems, and graceful failure rather than confident catastrophe. These are the unglamorous requirements that turn a marvel into infrastructure, and they are exactly where the current systems are weakest.
The distinction between capability and deployability is one the software world learned long ago and the intelligence world keeps having to relearn. A brilliant prototype that works in the lab is separated from a product people can depend on by an enormous, tedious span of engineering that has nothing to do with brilliance and everything to do with handling the cases the prototype never met. For a system meant to act generally in the world, that span is not merely enormous but possibly unbounded, because the world’s supply of edge cases is inexhaustible, and a general system, by definition, will eventually encounter all of them. The narrow tools that already run quietly inside critical infrastructure earned their trust by being predictable within tightly bounded domains. A system that can attempt anything forfeits exactly that boundedness, and with it the very property that made the narrow tools trustworthy in the first place.
The evidence is stark. On tasks that require actually operating in the world rather than answering questions about it, performance falls off a cliff: the best autonomous agents manage roughly half of what expert humans achieve on complex real research, and drop to something like a sixth of expert accuracy on genuinely messy real-world analysis, while robots still fail the large majority of ordinary household tasks and the gap between benchmark achievement and deployment reliability is openly acknowledged as one of the central problems the field has not solved. This is the definition of done that the whole catalog insists on: done means boring, a system so reliable and so predictable that you would trust it to run a Tuesday unsupervised, which is the opposite of a dazzling demo, and it is precisely the thing no one yet knows how to build, closely related to the hard integration problems that shadow every frontier medical technology, from the neural implants that must work flawlessly inside a living person onward. The demo is the easy part. The Tuesday is the mountain.
Artificial General Intelligence in 2026
The state of the field in 2026 is a vivid portrait of both truths at once, capability screaming upward while reliability and clarity lag behind. On the capability side, there is no plateau: scores on a benchmark of resolving real software issues leapt from sixty percent toward the human baseline to near-total in a single year, systems now match or exceed human experts on PhD-level science questions and competition mathematics, a frontier system won gold at the International Mathematical Olympiad, adoption reached the overwhelming majority of large organizations, and the technology diffused to more than half the population faster than the personal computer or the internet, all atop private investment measured in the hundreds of billions of dollars. By the standards of any earlier decade, this is general intelligence, arrived and deployed, as the field’s flagship annual accounting in the Stanford AI Index report on technical performance documents in exhaustive detail.
And on the other side of the same ledger sits the jagged, unreliable, unmeasurable reality: the clock-reading failures, the wide range of hallucination rates, the sharp collapse of accuracy on real-world tasks, the rising count of documented incidents, the benchmarks saturating faster than new ones can be built, and the falling transparency that makes vendor-reported scores an unreliable guide to real behavior. The definition wars, meanwhile, have descended into genuine absurdity, with the term artificial general intelligence appearing in corporate contracts pegged to dollar thresholds of profit, as if a mind could be defined by a revenue figure, a fittingly surreal endpoint for a target no one can specify. The honest live question in 2026 is therefore not how close we are to artificial general intelligence, which is unanswerable because the phrase names no measurable state, but whether the sprawling, jagged, general-ish capability we have actually built can be made reliable and verifiable enough to trust, a question whose answer depends on solving problems that more scale leaves entirely untouched, and that is increasingly entangled with the same chip-and-power geopolitics driving the global scramble over critical materials. The capability keeps climbing. The landing pad has still not been built.
The Moonshot With No Landing Pad
Strip artificial general intelligence to its foundation and the lesson generalizes past computing, because it is the strangest version of an error the whole catalog keeps making. Every other moonshot at least knows what it is trying to do, and fails on the unglamorous execution. This one cannot even state its goal, because it named itself after a property, generality, that it has largely achieved and that turned out not to be the point, while aiming away from the properties that are: reliability, verifiability, and the boring dependability that separates a demonstration from a foundation. It set as its finish line a horizon defined as the place it has not reached, guaranteeing that the line would recede at the speed of every advance, and it built no way to test whether it had crossed a line that was never drawn. These are not obstacles that more compute removes. They are the actual shape of the problem, and they were always the actual shape of the problem.
The realistic future, then, is the one already unfolding, and it is neither the utopia nor the apocalypse that dominates the discourse but something quieter and harder: the slow, unglamorous work of taking a staggeringly general and staggeringly unreliable capability and grinding it toward the narrow, verified, trustworthy reliability that real deployment demands, one jagged edge at a time, in specific domains where the answer key exists and the failure modes are survivable. The general mind that can be trusted to run a Tuesday unsupervised, across everything, recedes into an honest distance, guarded not by insufficient cleverness but by the absence of a definition, a test, and a solution to the reliability that was always the true wall. This is the entry in the catalog of civilization’s great technological moonshots where the honest move is to admit we cannot see the summit because we never agreed on where it was. We thought the hard part was making a machine general. It turns out we did that, almost by accident, and the machine is brilliant and erratic and cannot be trusted, and we still cannot say what it would mean to be done, or how we would know. The moonshot has no landing pad. It never did.
