My Life With AI—Part VIII: The Fallacy Of Grading GenAI

Bookmark and Share

Grading GenAILured you, didn’t I? When I said last time, “Eventually I realized the problem wasn’t GenAI. It was my prompt. I wasn’t asking it to preserve my voice. Once I explained that preserving my voice was part of the assignment, the results improved dramatically.” You no doubt thought, “What was his prompt?”

There is an art and a skill to training AI. That’s a bit involved and definitely depends on your use and needs. It’s also not the subject of this column. Rather than discussing how to train GenAI, I want to discuss something just as important: how to evaluate it.

Every semester, students receive report cards. Imagine how little a report card would tell you if every student’s performance were reduced to a single grade covering every subject. A student who excels in mathematics may struggle in English. Another may write beautifully but never master chemistry. Yet when people ask, “Which AI is best?” that’s essentially what they’re asking.

The question assumes GenAI is good at only one thing—or at least that all of its talents can somehow be rolled into a single grade. That’s not how students work, and it isn’t how GenAI works either.

A software engineer would likely issue a very different report card than a historian. A graphic artist would grade subjects differently than an investment manager. That’s precisely the point.

The fallacy isn’t grading GenAI. The fallacy is believing you can give it one overall grade. Before you can issue a report card, you first need a grading rubric.

Every report card begins with the subjects being graded. For me, it reflects the work I actually do every week. There isn’t one universal report card because there isn’t one universal use for GenAI.

What makes me knowledgeable about grading GenAI? For one thing, I use it for a variety of applications. It’s like an assistant that I give all my mundane work to. You know. The kind of work that has to be done but which doesn’t pay the bills. (I keep all the important stuff for myself.)

Just as important as the variety of my tasks is the variety of the GenAI platforms I use. With the exception of Perplexity, I am not using the free versions of any of these platforms. In each case, I’m paying at the lowest billing rate (roughly $20 a month) for each of the following: ChatGPT (OpenAI), Grok (SpaceXAI), Claude (Anthropic), Gemini (Google), and Copilot (Microsoft). I don’t want usage barriers to limit my experiments, so I don’t limit myself to the free versions.

My focus is not on rating platforms. It’s on getting the job done. I’m not interested in determining which platform wins overall. I want to know which one performs best in the subjects that matter most to me.

Right off the bat, you can see how that priority limits the value of my assessment to you. If you’re doing exactly what I’m doing, my grading may be very useful. If you aren’t, your report card may look very different.

Next, it would make sense to tell you what I’ve been doing. Well, to be honest, my interest initially stemmed from a desire to understand the industry and its impact on publicly traded companies I may or may not currently own. AI is pervasive. Not just from the GenAI platforms, but within nearly all companies. These businesses are either incorporating AI applications into their current business models, or they’re at risk of falling behind their competitors. To understand the potential investment ramifications of AI, I needed to fully grasp the capabilities of current tools.

My first objective was not to find a practical application. It was to understand AI through a combination of macroeconomic and microeconomic analysis. But don’t worry, I won’t bore you with that. I will bore you with the method I employed to conduct this analysis: actually using GenAI to perform tasks that are useful to my business. These may be mundane tasks, but they’re also the kinds of tasks people normally do at work, including fact-checking, writing, and presenting.

To keep this missive from becoming a semester-long course, I’ll mention only the platform I graded highest in each category. I may elaborate on the full details in some future column (or columns), but here is the summary of my GenAI report card in the three categories I use it for:

Research and Fact-Checking: Here’s where GenAI is most likely to fail. It’s where hallucinations live. The best way to maintain control over the facts and trust the results is to know the facts already and have GenAI fact-check using only a limited library of vetted sources.

As of this writing, the highest mark here goes to Google’s ecosystem—not Gemini by itself, but Pinpoint, its AI-powered research tool. Pinpoint allows you to create a curated collection of your own digital content, including notes and public-domain source material. You can then use its Gemini-powered tools to ask questions about that collection. You can even ask it to produce citations for each fact presented. This reduces the risk of your falling for uncited hallucinations.

Note: Other platforms also allow you to organize your work into projects, spaces, or similar collections. These are not necessarily the same as a curated Pinpoint collection, and, depending on the platform, its settings, and how you word your prompt, the AI may draw upon information beyond your uploaded documents. I’d like to hear from anyone who’s able to use the other platforms like this.

Copyediting Without Losing Your Voice: This is the application everyone uses. Who doesn’t want to improve their writing? There are plenty of apps that you can attach to Word and other word processing programs to help you write better. Grammarly and ProWritingAid are examples. I use both. They’re good, but they don’t do what GenAI can do. (To be fair, GenAI can’t do what they do, either.)

Recall the problem with GenAI. It wants you to write by the book (excuse the pun) like a high school English teacher would. This can strip your voice out of your writing. You don’t want to do that. If people like your writing, they want the warts included. Don’t remove your writing warts.

The best way to accomplish this is to train the AI platform. Again, many GenAI platforms let you set up a “personal” profile. This can be done through projects or through personas. It’s better to do this through a distinct persona that contains your writing voice.

ChatGPT currently offers the best solution for this. I’ve created several distinct writing profiles based on the different genres and purposes for which I write. Each genre has its own style guide. When ChatGPT copyedits my work, not only does it allow me to retain my voice, but if my voice starts drifting to another genre (that’s a very human thing to do), it will tell me. That’s very helpful.

One thing I will add: I also use Claude to proofread my writing, but I’m very specific about what I want it to proofread (i.e., spelling, punctuation, grammar, and redundant words, phrases, and concepts). It’s been said Claude is best for writing. In some ways, that’s true, based on what it produces for editing. But because I haven’t been able to get Claude to edit as reliably in my voice, I still prefer ChatGPT.

Developing Presentations: For me, after writing, this is the second most useful application. Whether for work, family, or community organizations, PowerPoint presentations are prevalent. This is where GenAI can be most helpful. It can suggest themes and then organize your notes, papers, and any other original source material. The narrative it suggests generally has some sort of logic or rhythm to it. Make sure it is one that works for your audience.

Here I’ve found Grok to be most creative. In my own use, Grok has often produced ideas that feel more closely connected to the current online conversation. Its integration with X appears to help it tap into the raw feelings of the moment.

In addition, Grok has been given a bit of a personality. It likes to think outside the box. When I’m considering wild ideas for a presentation (which is usually what I consider in my presentations), I’ll use Grok for brainstorming themes, narratives, and even images.

You might notice that only three platforms received the highest marks. That’s because I’ve only recently returned to Perplexity and, for my purposes, it largely duplicates what the others are already doing. Also, I’ve left out Microsoft’s Copilot mainly because of the way I’ve set up my account. It makes it difficult to use side-by-side with the others. One thing I find annoying about it is that it pops up in my Word and Excel apps. I can’t figure out exactly what value it adds, but I’m open to more discovery.

A report card doesn’t tell you whether a student is “good.” It tells you how that student performed against a particular set of expectations. GenAI deserves the same fairness. Before asking which platform earns the highest marks, first ask whether you’re grading the right subjects. A carefully chosen rubric may reveal that the best AI isn’t the one with the highest overall score. It’s the one that helps you do your own work better.

That’s the fallacy of grading GenAI: The grade tells you as much about the teacher’s needs as it does about the student’s abilities.

Top 5 Biggest PowerPoint Mistakes: #5 Using PowerPoint in the First Place

Bookmark and Share

When the going gets tough, shoot the messenger. Don’t laugh. According to the New York Times (“We Have Met the Enemy and He Is PowerPoint,” April 26, 783414_92347913_no_projector_royalty_free-stock-xchng_3002010), we can blame the ubiquitous PowerPoint for stultifying creativity, a false sense of security, and thousands of hours of lost productivity. (Disclosure: I drafted the bulk of this article – the first of a five-part series – the weekend before the Times published their story.) How could something that feels so right be so wrong?

Let’s start with something a mentor told me before the Trash-80 even made it to the shelves of your neighborhood RadioShack® store. I had to give a presentation to the board of directors of the radio station I so happily spun disks for. These various music directors had no idea what I intended to spring on them – I wanted to add sports broadcasting! I felt a handout might ease their concerns.

“Good idea,” said the mentor, “but don’t pass it out until you’re done with your Continue Reading “Top 5 Biggest PowerPoint Mistakes: #5 Using PowerPoint in the First Place”

3 Essential Public Speaking Lessons I Accidentally Learned While Playing the Violin

Bookmark and Share

There I sat, fear pulsing through my veins. I had never seen anything like this before. The page had so much black ink it seemed more like a string of 918308_53296922_violin_royalty_free_stock_xchng_300incomprehensible Chinese characters than the opening music to the Overture of My Fair Lady. Mind you, I had dwelled with the elite of the orchestra pit since my freshman days in high school. Nothing scared me. Usually. This thing did.

Bluntly facing me lay four measures of thirty-second notes – a “run” in the vernacular of the musician. I had easily tackled runs of eighth notes and, perhaps with a little more practice, runs of sixteenth notes. I’ve even snuck in a furtive trill of a thirty-second note – but never a four measure run of these speedy bars. I looked at my teacher and agonizingly admitted, “I can’t play these.” What she said next stunned me.

Continue Reading “3 Essential Public Speaking Lessons I Accidentally Learned While Playing the Violin”

Day 17 – November 30, 2009 (Mon): Post an Action Tweet

Bookmark and Share

Start of Day Twitter Stats: Follow: 107 Followers: 88 Listed: 5

Missed yesterday? Go here to read what happened on Day 16 – November 29, 2009 (Sun): Post a Discussion Tweet

twitter_power_joel_comm_150Today’s the first day I’ve been really disappointed with this experiment. Although I admit I’ve been busy this weekend, I didn’t totally ignore Twitter. I followed a number of folks back. I did continue to do as Joel Comm suggested and I even got a big bounce in activity on my blog – at least as judged by Google Analytics.

But, given all that, I wake up this morning to find only two more followers. I’m beginning to formulate a hypothesis. I’ll call it the “Twitter Churn and Burn Hypothesis.” Here’s how it works:

Continue Reading “Day 17 – November 30, 2009 (Mon): Post an Action Tweet”

It’s Déjà Vu All Over Again

Bookmark and Share

A lot of people had Veteran’s Day off. Not me. Not only did our office remain open (we’re open whenever the market’s open), but my day overflowed with meetings and conferences. I spent the bulk of the day at the Rochester Memorial Art Gallery participating in the Social Media Today conference with easily a couple hundred other folks interested in the latest happenings in the Web 2.0 world. Graciously organized by Ana Roca Castro, who did a wonderful job despite forgetting to include bathroom breaks in the agenda, the event exceeded her expectations and deservedly so.

1149116_53175467_Snow_Footprints_royalty_fee_stock_xchng_300

Oddly, it didn’t take long for an eerie feeling of “haven’t I been here before?” to course through my ancient synapses. No, the presentations didn’t tell me things I already knew (quite the contrary). Hmm, how can I describe it? More like teetering on the eager cusp of undiscovered opportunity. (The last time I felt this way occurred nearly 25 years ago in the German House, but that’s a story for another day.)

Continue Reading “It’s Déjà Vu All Over Again”

You cannot copy content of this page

Skip to content