It's pretty close to how we measure IQ. The standard test is basically a series of spatial puzzles.
I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we saw the Turing test as the proxy for AGI, but then early LLMs could easily pass for a human in a casual conversation while clearly not matching human performance on most other tasks.
Since then, every benchmark we come up with, it turns out that an LLM can be fine-tuned to solve it while still clearly lacking something. They make very non-human mistakes, are easily tricked because they have a pretty tenuous grasp of reality, etc. But I think this just shows that AGI is a meaningless marketing term. We could as well be arguing if they have souls.
When I was 18, my high school girlfriend took me to the local Mensa chapter’s New Year’s party because her mother was a member and she was used to hanging out there.
It was a useful lesson that whatever IQ tests measure, it is completely devoid of value or interest to me.
At the risk of sounding like one of those people at Mensa that annoyed you...
The people at Mensa aren't a valid sample of people who score high on IQ tests, because there is a such a strong selection effect for people with certain personality traits, such as wanting to join a club based on your IQ.
By looking at interactions of people in Mensa I feel like 2% of top IQ is actually too low bar to notice anything particular about "high" IQ people. We are still pretty random in our interests and performance.
Is it a serious question? IQ tries to quantify the positive correlation between the results of all intellectual tasks a person takes (AKA positive manifold).
I wanted to highlight that we're actually looking at a trio of concepts: intelligence, IQ, and value. Strongly correlated concepts, yes, but also meaningfully distinct!
Whatever the people who came up with IQ intended isn’t really relevant to the question of whether IQ measures intellect. At best is is loosely correlated. Very loosely.
IQ strongly correlates with many real world outcomes that people tend to associate with being more "intellectual".
IQ has precise definition. It's what the tests measure. Intellect has only fuzzy handwavey definition. Correlation between IQ and intellect is about as loose as the definition of the intellect. The way people make the correlation even looser is by defining intellect in even more fuzzy and nebulous manner.
I suspect my IQ isn’t quite high enough to join Mensa but I’ve always flirted with the idea of trying to join and getting in just to see what a group of Mensa people are like.
We are basically random. 2% of population by top IQ test result has pretty much very similar distribution in all aspects to 100% of the population. It's not a high bar to be fastest processing one among 50 people.
If you calculate this properly, a Mensa member has about a cointoss chance to be the "smartest" among random group of 35 people (not 50).
Given that they are probably rarely in groups of purely random people, if you are Mensa member and you are in a group of dozen people (for example in professional setting) you'd probably have a 50/50 chance of not being the one with the highest IQ there.
What should I research? IQ follows normal distribution across entire population. I guess if you'd built your heterogenous population out of mental patients and university professors together you'd get bimodal distribution, but such heterogenous populations don't spontaneously pop up in your life very often.
IQ tests are incredibly good at what they're designed for, which is discriminating relatively higher intelligence humans from lower intelligence humans. Also discriminating within a single human - they are routinely and reliably used to track cognitive decline.
For these purposes they are highly reliable (repeatable, internally consistent) and valid (correlate with ~everything to about the degree one would reasonably expect).
They were never designed for machines or non-human animals.
Nor were they designed for rare ranges of intelligence - these are by definition hard to create tests for, since it's hard to gather the sample sizes you need. So they work well for the middle ~98% of humans but can't discriminate well among the most profoundly intellectually disabled nor among true geniuses.
It might, but if there's one thing that hasn't changed since 2022, it's that the models tend to ace the tasks where you have gobs of training data and where verification loops are fast and cheap... and they are not nearly as amazing elsewhere. If it's close enough, they can generalize, e.g. translate one programming language to another. But there's a pretty steep cliff past a certain distance.
Case in point: you had hundreds of millions of JPEGs to vacuum up and bitmap image generation is amazing. But if you ask them to recreate the same scene as vector art, they will struggle to generate a decent SVG. Like, kindergarten-style pelicans on bicycles are the state of the art. It should generalize seamlessly, but somehow, doesn't?
I think it will happen, just like self-driving cars are happening, but it will probably be a slow process.
Nit: Turing’s actual imitation game is a party game (like Werewolf/Mafia) and nobody’s even trying to win at that. The LLM’s will just tell you they’re an AI.
This is an artifact of how we deliberately craft these models though. We could just as easily fine tune a model that will believe it is not an AI or will attempt to deceive users asking about it
There's a lot more to it. For example, its writing style would also have to improve so it doesn't immediately give itself away.
Also, the skill of the human opponents matters. You'd want to test it against people who have practiced playing the game. Otherwise, it's like the difference between building a chess bot that can win against random undergrads who don't normally play, versus winning against grandmasters. And it's not like there's a pool of skilled human players of the imitation game.
Their input and output interfaces are too different from human's and they're not nearly as smart to take our IQ tests, but both dolphins and octopuses can solve complex puzzles tailored for their environment. Those puzzles are the whole reason scientists know that dolphins and octopuses are more intelligent than other animals.
But we do know they're "intelligent" and also smart in an important capacity. So how gives we don't measure them by IQ? Because the IQ is not a good measure of intelligence or smarts.
> No, it is just because they have difficulties at the bench.
I'll put it in another way. A "gifted kid" can be measured incredibly well on an IQ test, but fail miserably at incredibly normal but very difficult tasks such as consoling someone for their loss and managing family crisis. This is a clear example where an IQ measure doesn't translate to a person being capable of meaningfully changing their environments for good which is one way we define intelligence.
On the other hand saying "the gifted person is highly intelligent/smart just not good at some things" really diminishes the other tasks, because they really are very difficult tasks but are not measured by an IQ test.
I don’t follow your logic. “The gifted person is highly intelligent, just not good at deadlifting 500kg” does not diminish the 500kg deadlift and I wouldn’t expect an IQ test to measure it.
Deadlifting 500kg is not a form of intelligence (or it could be under certain scenarios). I carefully chose specific tasks for my comment, because those tasks do reflect a kind of intelligence that's not measured by IQ.
I think the key here is to be wary of measurements that promise to capture the whole of what we consider intelligence (ie what people think of with IQs).
That would be odd framing: it's a good test for a form of intelligence and other forms of intelligence still require good tests. It remains a good metric - for its specific thing; those who believe it to be the whole metric are naïve. It is almost necessary though not really sufficient.
I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we saw the Turing test as the proxy for AGI, but then early LLMs could easily pass for a human in a casual conversation while clearly not matching human performance on most other tasks.
Since then, every benchmark we come up with, it turns out that an LLM can be fine-tuned to solve it while still clearly lacking something. They make very non-human mistakes, are easily tricked because they have a pretty tenuous grasp of reality, etc. But I think this just shows that AGI is a meaningless marketing term. We could as well be arguing if they have souls.