My Life With AI—Part VIII: The Fallacy Of Grading GenAI

Bookmark and Share

Grading GenAILured you, didn’t I? When I said last time, “Eventually I realized the problem wasn’t GenAI. It was my prompt. I wasn’t asking it to preserve my voice. Once I explained that preserving my voice was part of the assignment, the results improved dramatically.” You no doubt thought, “What was his prompt?”

There is an art and a skill to training AI. That’s a bit involved and definitely depends on your use and needs. It’s also not the subject of this column. Rather than discussing how to train GenAI, I want to discuss something just as important: how to evaluate it.

Every semester, students receive report cards. Imagine how little a report card would tell you if every student’s performance were reduced to a single grade covering every subject. A student who excels in mathematics may struggle in English. Another may write beautifully but never master chemistry. Yet when people ask, “Which AI is best?” that’s essentially what they’re asking.

The question assumes GenAI is good at only one thing—or at least that all of its talents can somehow be rolled into a single grade. That’s not how students work, and it isn’t how GenAI works either.

A software engineer would likely issue a very different report card than a historian. A graphic artist would grade subjects differently than an investment manager. That’s precisely the point.

The fallacy isn’t grading GenAI. The fallacy is believing you can give it one overall grade. Before you can issue a report card, you first need a grading rubric.

Every report card begins with the subjects being graded. For me, it reflects the work I actually do every week. There isn’t one universal report card because there isn’t one universal use for GenAI.

What makes me knowledgeable about grading GenAI? For one thing, I use it for a variety of applications. It’s like an assistant that I give all my mundane work to. You know. The kind of work that has to be done but which doesn’t pay the bills. (I keep all the important stuff for myself.)

Just as important as the variety of my tasks is the variety of the GenAI platforms I use. With the exception of Perplexity, I am not using the free versions of any of these platforms. In each case, I’m paying at the lowest billing rate (roughly $20 a month) for each of the following: ChatGPT (OpenAI), Grok (SpaceXAI), Claude (Anthropic), Gemini (Google), and Copilot (Microsoft). I don’t want usage barriers to limit my experiments, so I don’t limit myself to the free versions.

My focus is not on rating platforms. It’s on getting the job done. I’m not interested in determining which platform wins overall. I want to know which one performs best in the subjects that matter most to me.

Right off the bat, you can see how that priority limits the value of my assessment to you. If you’re doing exactly what I’m doing, my grading may be very useful. If you aren’t, your report card may look very different.

Next, it would make sense to tell you what I’ve been doing. Well, to be honest, my interest initially stemmed from a desire to understand the industry and its impact on publicly traded companies I may or may not currently own. AI is pervasive. Not just from the GenAI platforms, but within nearly all companies. These businesses are either incorporating AI applications into their current business models, or they’re at risk of falling behind their competitors. To understand the potential investment ramifications of AI, I needed to fully grasp the capabilities of current tools.

My first objective was not to find a practical application. It was to understand AI through a combination of macroeconomic and microeconomic analysis. But don’t worry, I won’t bore you with that. I will bore you with the method I employed to conduct this analysis: actually using GenAI to perform tasks that are useful to my business. These may be mundane tasks, but they’re also the kinds of tasks people normally do at work, including fact-checking, writing, and presenting.

To keep this missive from becoming a semester-long course, I’ll mention only the platform I graded highest in each category. I may elaborate on the full details in some future column (or columns), but here is the summary of my GenAI report card in the three categories I use it for:

Research and Fact-Checking: Here’s where GenAI is most likely to fail. It’s where hallucinations live. The best way to maintain control over the facts and trust the results is to know the facts already and have GenAI fact-check using only a limited library of vetted sources.

As of this writing, the highest mark here goes to Google’s ecosystem—not Gemini by itself, but Pinpoint, its AI-powered research tool. Pinpoint allows you to create a curated collection of your own digital content, including notes and public-domain source material. You can then use its Gemini-powered tools to ask questions about that collection. You can even ask it to produce citations for each fact presented. This reduces the risk of your falling for uncited hallucinations.

Note: Other platforms also allow you to organize your work into projects, spaces, or similar collections. These are not necessarily the same as a curated Pinpoint collection, and, depending on the platform, its settings, and how you word your prompt, the AI may draw upon information beyond your uploaded documents. I’d like to hear from anyone who’s able to use the other platforms like this.

Copyediting Without Losing Your Voice: This is the application everyone uses. Who doesn’t want to improve their writing? There are plenty of apps that you can attach to Word and other word processing programs to help you write better. Grammarly and ProWritingAid are examples. I use both. They’re good, but they don’t do what GenAI can do. (To be fair, GenAI can’t do what they do, either.)

Recall the problem with GenAI. It wants you to write by the book (excuse the pun) like a high school English teacher would. This can strip your voice out of your writing. You don’t want to do that. If people like your writing, they want the warts included. Don’t remove your writing warts.

The best way to accomplish this is to train the AI platform. Again, many GenAI platforms let you set up a “personal” profile. This can be done through projects or through personas. It’s better to do this through a distinct persona that contains your writing voice.

ChatGPT currently offers the best solution for this. I’ve created several distinct writing profiles based on the different genres and purposes for which I write. Each genre has its own style guide. When ChatGPT copyedits my work, not only does it allow me to retain my voice, but if my voice starts drifting to another genre (that’s a very human thing to do), it will tell me. That’s very helpful.

One thing I will add: I also use Claude to proofread my writing, but I’m very specific about what I want it to proofread (i.e., spelling, punctuation, grammar, and redundant words, phrases, and concepts). It’s been said Claude is best for writing. In some ways, that’s true, based on what it produces for editing. But because I haven’t been able to get Claude to edit as reliably in my voice, I still prefer ChatGPT.

Developing Presentations: For me, after writing, this is the second most useful application. Whether for work, family, or community organizations, PowerPoint presentations are prevalent. This is where GenAI can be most helpful. It can suggest themes and then organize your notes, papers, and any other original source material. The narrative it suggests generally has some sort of logic or rhythm to it. Make sure it is one that works for your audience.

Here I’ve found Grok to be most creative. In my own use, Grok has often produced ideas that feel more closely connected to the current online conversation. Its integration with X appears to help it tap into the raw feelings of the moment.

In addition, Grok has been given a bit of a personality. It likes to think outside the box. When I’m considering wild ideas for a presentation (which is usually what I consider in my presentations), I’ll use Grok for brainstorming themes, narratives, and even images.

You might notice that only three platforms received the highest marks. That’s because I’ve only recently returned to Perplexity and, for my purposes, it largely duplicates what the others are already doing. Also, I’ve left out Microsoft’s Copilot mainly because of the way I’ve set up my account. It makes it difficult to use side-by-side with the others. One thing I find annoying about it is that it pops up in my Word and Excel apps. I can’t figure out exactly what value it adds, but I’m open to more discovery.

A report card doesn’t tell you whether a student is “good.” It tells you how that student performed against a particular set of expectations. GenAI deserves the same fairness. Before asking which platform earns the highest marks, first ask whether you’re grading the right subjects. A carefully chosen rubric may reveal that the best AI isn’t the one with the highest overall score. It’s the one that helps you do your own work better.

That’s the fallacy of grading GenAI: The grade tells you as much about the teacher’s needs as it does about the student’s abilities.

You cannot copy content of this page

Skip to content