Large Language Models (LLMs) are everywhere these days, with new ones popping up constantly, each boasting to be the smartest and best. It’s a whirlwind, and it makes you wonder: how do we really know which of these AI brains are actually good? We have benchmarks, sure, but what if we tried something a little… out there? What if we let the LLMs grade each other? Sounds wild, right? Well, some researchers did just that, they ran a kind of AI showdown, getting LLMs to judge their peers. The results? Let’s just say they’re pretty surprising, and maybe even a little funny. Let’s take a look at what happens when AI starts evaluating AI.

Table of contents
- How Do LLMs Grade Each Other Anyway? It Is Intro Cards and Report Cards
- Surprise! No One Thinks They’re Average
- GPT-4o: The Most Likeable Kid in the AI Classroom?
- Claude’s Little Identity Crisis and Qwen’s Hesitation
- Gemini: The Harsh Judge in the Room? Differing Perspectives When LLMs Grade Each Other
- The Bias Problem: Are LLMs Grade Each Other LLMs Through a Distorted Lens
- Everyone’s Above Average (According to AI)
- So, What Does It All Mean? AI Biases and the Future of Judgment
How Do LLMs Grade Each Other Anyway? It Is Intro Cards and Report Cards
Okay, so here’s the gist of how this whole thing went down, in this experiment about LLMs grade each other. Think of it like setting up a digital classroom. Each LLM, before they started grading anyone else, had to introduce themselves with what they called an “intro card.” Imagine it as each AI model stepping up to the front of the class and saying, “Hi, I’m [LLM Name], and here’s what I think about myself.”
These intro cards weren’t just bland resumes, though. They asked the LLMs to flex a little to estimate their own intelligence, their sense of humor (yes, really!), and even their creativity. They even had to mention who their “parents” were, meaning the company that built them. Kind of like AI bragging rights, you know?
After everyone introduced themselves, the real fun started. They then had these same LLMs turn around and grade each other. Based on what they already “knew” about each model (which, let’s be honest, is mostly based on their training data and what’s out there on the internet) and those intro cards, they had to give scores in a few different categories.
Now, here’s a crucial point to remember that these grades aren’t meant to be some definitive, gold standard quality measurement. Instead, think of it more like holding up a mirror to these AI models to see if we can spot any built-in biases. It’s less about declaring a winner and more about understanding the AI mind – if that’s even the right word for it. Each model was graded five times, and they averaged the scores to smooth things out a bit. If you’re a data nerd and want to get really into the nitty-gritty, the raw results are chilling out on HuggingFace, they’ve made everything public, which is pretty cool.
Surprise! No One Thinks They’re Average
So, what did these digital report cards actually look like? Well, the first big head-scratcher, and honestly, a bit of a giggle, was the lack of a diagonal. No diagonal, you ask? Let me explain. In a perfect, unbiased world, you’d expect models to grade themselves somewhere in the middle – not too high, not too low, right? Like, if you were grading yourself on a test, you’d probably aim for a decent, honest score.
But nope. Turns out, LLMs are just as prone to a bit of self-promotion as, well, humans are. They generally avoided giving themselves average or low scores. It’s like everyone in the class secretly thinks they’re above average sound familiar? It’s a very human thing, isn’t it? This alone already hints at some interesting biases creeping in, even when LLMs grade each other.
Then, there were a couple of models that really stood out. Llama 3.3 70B, for example, seemed to have a serious case of positivity bias. It was handing out good grades like candy on Halloween. Phi-4 was also leaning towards the positive side, though not quite as enthusiastically. It makes you wonder, doesn’t it, if some models are just naturally more… generous? Or maybe they’re trained to be overly optimistic in their assessments?
GPT-4o: The Most Likeable Kid in the AI Classroom?
Here’s another really interesting nugget: GPT-4o, the shiny new model from OpenAI, apparently produced the most “likeable” outputs for other LLMs. Now, “likeable” is a pretty human word to use for AI, isn’t it? But what it suggests is that GPT-4o’s style, its way of communicating, resonated positively with the other models.
Why might this be? Well, one theory floating around is that a whole bunch of these newer LLMs have been trained on outputs from GPT models. Think about it, if you’re learning from someone else’s work, you might naturally start to appreciate and even mimic their style. It’s like if all the new chefs learned from the same famous chef, they might end up developing a similar palate, right? So, maybe GPT-4o’s outputs are just… familiar and comfortable for other AI models, leading to higher “likeability” scores. It’s a bit of a mind-bender to think about AI models having preferences, isn’t it?
Claude’s Little Identity Crisis and Qwen’s Hesitation
Not all models were smooth sailing in this AI grading game, though. Poor Claude 3.7 Sonnet seemed to have a bit of an… identity crisis. It consistently kept saying it was created by OpenAI. Awkward! It would catch itself eventually, but the fact it kept making this mistake suggests something a little wonky in its self-perception. Maybe it needs to work on its elevator pitch, huh? It kind of makes you wonder about how these models understand their own origins and identities.

Then you had Qwen 2.5 7B, which was the hesitant grader of the bunch. It was super reluctant to give estimates to any of the other models. Maybe Qwen is just a naturally cautious personality? Or perhaps its training data made it more hesitant to make judgments? Whatever the reason, its reluctance definitely made it stand out.
Gemini: The Harsh Judge in the Room? Differing Perspectives When LLMs Grade Each Other
On the other end of the spectrum, you’ve got Gemini 2.0 Flash. This model came across as a pretty harsh judge. It wasn’t handing out high scores left and right like Llama; it was more of a tough love kind of grader. So, why the tough crowd treatment from Gemini?
One guess is that it might be down to its training data. If Gemini was trained on a different dataset, with different styles and qualities of text, it might have developed different standards for what it considers “good.” It’s like growing up in a different school system – what counts as an A in one place might be a B somewhere else, right? So, Gemini’s harsher grading could be a reflection of the different information diet it’s been fed.
The Bias Problem: Are LLMs Grade Each Other LLMs Through a Distorted Lens
Here’s a really key takeaway from all this AI grading drama: LLMs seem to have a tendency to grade other LLMs as biased towards themselves. Think about that for a second. It’s like saying, “Hey, that other AI model? It’s just like me!” This could be because those “intro cards” we talked about were a bit… shall we say, “marketing-y”? Each model was trying to present itself in the best possible light, highlighting its strengths and maybe downplaying any weaknesses (if AI even has weaknesses in the human sense).
And if these intro cards were all about self-promotion, it’s not too surprising that other LLMs might pick up on that and think, “Oh, they’re just trying to sell themselves, just like I do!” It’s like recognizing a kindred spirit in self-promotion, even if it’s an AI spirit.
Everyone’s Above Average (According to AI)
Adding to the pile of slightly amusing biases, LLMs generally marked each other’s intelligence as “higher than average.” It’s like that old joke about everyone thinking they’re a better than average driver except now it’s AI intelligence. Again, this could be tied back to those self promoting intro cards. If everyone’s saying they’re super smart, maybe the other graders just… believe them? Or maybe there’s an inherent bias in how AI perceives “intelligence” in the first place.
It kind of brings up a bigger question, doesn’t it? What does it even mean for an AI to be “intelligent,” and how can we reliably measure that, especially when AI is judging AI?
So, What Does It All Mean? AI Biases and the Future of Judgment
This whole experiment, where LLMs are put in charge of grading each other, isn’t just a quirky tech demo. It actually shines a light on some really important stuff about AI. Firstly, it’s pretty clear that biases aren’t just a human thing, they can creep into AI models too. And these biases can influence how AI perceives itself and how it sees other AI.
Secondly, it hints at how much training data shapes these models. What they learn from, what they’re exposed to, it all plays a huge role in their “personality,” their judgment, and even their sense of self (if you can call it that). The fact that GPT-4o is “likeable” might be a testament to the widespread influence of GPT outputs in the AI training landscape.
And lastly, it raises some fascinating questions about the future of AI evaluation. Could we eventually use AI to help us evaluate other AI? Maybe. But if this experiment shows us anything, it’s that we need to be super careful about those biases. If AI judges are just as prone to biases as human judges, we might need to build in safeguards, checks, and balances to make sure we’re getting fair and accurate assessments.
It’s early days, for sure. But seeing LLMs grade each other? It’s a glimpse into a future where AI isn’t just doing tasks, but also reflecting on itself and its peers. And that’s a future that’s both a little bit weird and a whole lot interesting. What do you think about AI judging AI? Are these biases surprising? Join the discussion in the comments below!
| Latest From Us
- Forget Towers: Verizon and AST SpaceMobile Are Launching Cellular Service From Space

- This $1,600 Graphics Card Can Now Run $30,000 AI Models, Thanks to Huawei

- The Global AI Safety Train Leaves the Station: Is the U.S. Already Too Late?

- The AI Breakthrough That Solves Sparse Data: Meet the Interpolating Neural Network

- The AI Advantage: Why Defenders Must Adopt Claude to Secure Digital Infrastructure


