10 Comments
User's avatar
Alex Tolley's avatar

If only lawyers would use Dan Davies' method to check that the genAI spitting out a brief to check the output. Maybe a few more expensive fines, perhaps even a disbarment, might change behaviors. In those sort of cases, I am not sure if a better way is to ask the AI to provide cases with references relevant to the case, with a summary of why each case law might be used. Then check everything, use the cases as narrow documents in a RAG, and ask the AI to create a brief only with those curated cases. Then read the brief carefully to make sure it is still correct.

GenAIs should start getting better at retrieving relevant references than the enshittified Google search that is now worse than it was. Surely Google understands that offering a free Gemini AI search summary is undermining their enshittified, ad-supported search results?

It may be somewhat recursive, but can genAI be asked to provide a good prompt for a desired output?

It seems to me that the best way to use genAI is like the way one uses a spreadsheet or code, use it to test out ideas quickly. A spreadsheet removes the slow, repetitive calculation task and quickly shows the result. GenAI should be able to quickly test out how to solve a task, which the user then curates/refines, to produce a good task output. Humans and AIs working together, with the AI acting like a fast go-fer assistant,

I don't know where AI goes next, other than to respond more effectively to prompts. I cannot help but feel that Asimov got there with controlling robots many decades ago when he wrote that the Spacers, with many robots to control, were far better at voice commands than the Earthers, who had little, if any, exposure to robots. We may just learn how best to interact with AIs to get the desired response. It would certainly be nice to use voice input to state whether accuracy or creativity was important in the response, perhaps using a human role as the initial prompt, just as we sort of use "You are a knowledgeable assistant on this subject [X]". I would prefer a catalog of domain specialist AIs, rather than trying to tailor the responses of a general AI. I also think that domain specialists may be able to sell their expertise as an off-the-shelf AI. A BradBot that can answer questions on economic history for a trifling amount might be very useful, with reviewers providing star ratings as feedback. I also would like to stage debates between domain expert AIs to determine which argument[s] are a better way forward, and to expose flawed logic. IOW, different AIs rather than a "mixture of experts" model.

George Kappus's avatar

Before I rely on an AI site for any task, I think it’s useful to try an experiment to see how it does on an obscure subject where you already know the answer to the query you’re going to post. An example: I spend a fair amount of time birding in Colombia. When I first tried Claude I posed it a query: What is the most critically endangered bird species in Colombia? I happened to know the answer to the question because I bid in Colombia with a scientist who is studying the species and working on its conservation. The answer is Antioquia Brushfinch, of which there are approximately 125 surviving adult individuals in the world. The identity of the brushfinch is not exactly a secret; there are lots of articles on its status. Yet Claude came back with an answer that just listed a set of widely known, somewhat endangered species. My follow up and their response went like this (paraphrasing : Hey, Bozo, what about Antioquia Brushfinch? Oh, sorry sir, we weren’t trying to answer the question you asked, just give some examples (all of which have populations over 125); after all who even really knows which species is most endangered. But, come to think about it, the brushfinch is pretty endangered and here are some more random examples. The moral of the story: Don’t trust AI unless they cite authority that you can confirm actually exists, is published in a reputable publication, actually says what they say it says and hasn’t been seriously called into question. In short, AI is only useful to the extent that it gets you quickly to some resource that does answer your question. (I’m assessing it only as a research to, not as a tool for writing code, for example.)

glc's avatar

Here's a nice example.

https://amandaguinzburg.substack.com/p/diabolus-ex-machina

It will "lie" on every possible occasion - but it can't lie, since it only generates word (token) strings and does not, itself, assign them any meaning. The lie only happens when you read them and treat them as meaningful. Death of the author writ large.

George Kappus's avatar

In my example, there was nothing untrue in Claude’s original response. It just was not responsive to my query and completely missed the readily knowable best correct answer to my actual question.

George Kappus's avatar

Oops. This was intended as a reply to Kent’s post.

Kaleberg's avatar

Two thoughts about AI:

1) The whole point of a natural language interface is to avoid the need for prompt engineering. Consider what happened with Apple's English like Applescript. It looks like natural language, perhaps with a bit of stilted jargon. Unfortunately, it isn't quite natural language, so it is surprisingly tricky to use. There are statements it understands wonderfully and there are statements where it's wretchedly obtuse. People still write code using Applescript, but it's a niche language skill.

There's a long history of almost natural languages being seen as removing the need for expertise. I worked on a natural language database query system back in the very early 1970s. We were porting it from an SDS 940, read demotic papyrus, to a PDP-10, read parchment. Buckminster Fuller liked it. That old. There have been many such systems since. They look straightforward, but they lose favor when one needs to RTFM or repeatedly try variations and refinements to get one's work done.

There's an uncanny valley that AI could get stuck in, good enough to convince you that it's as good as natural language but not as good as talking to a moderately intelligent native language speaker. It's still early in the game, so we'll have to see.

2) The other point involves monetization. The big money is likely in advertising. That was where the big money was in search. That is where the big money was in social media. There are billions of dollars to be made by doing AI product placement. The first type will be context based. A query about something will serve up relevant ads as part of the answer. The next type will be user surveillance based. A query, especially at the lower tiers, will serve up relevant ads as part of the conversation.

With advertising in place, there's even more money to be made if AI gives bad answers. If the user has to iterate with repeated prompt engineering, that means more opportunities to suggests products relevant or not. It's a horrifying thought, but that's what happened with search engines. Once they were monetized with advertising, bad search paid better than good search. Even if there are people who find AI useful and are willing to pay for the better versions, the free tier that most people are going to be exposed to is going to be a loaded with advertising and provide poor answers at best.

----

I wish I could be more optimistic. My own explorations with AI haven't been particularly compelling. Mathematically, there has to be a median household discretionary income in the United States, but AI doesn't know what it is. AI believes that Korkor Lodge in Tigray is still open, but a quick search reveals that it was destroyed in the recent war. AI may actually be useful, but I am still not convinced I should trust it.

Kent's avatar

In my mind, LLM's are running thousands of Cartesian joins of the Internet with itself. This is no doubt inaccurate, but as a metaphor, there are going to be some cells with sparse data, and many cells with data derived from low quality sites. The next word in the LLM sentence can be the equivalent of the Family Feud game: "let's see what the answer is based on our survey."

It would be helpful if LLM's returned confidence intervals on their answers, because they sound so damn confident that if you don't already know the answer, then you'd be sure they've got the goods. LLM's need to communicate those three little beautiful words, "I don't know", or at least, "I'm not sure."

Matt Curtis's avatar

I think Google Search is a better analogy than the calculator. At least when it comes to rules for students. Schools had to adjust quite a bit to Google Search + Wikipedia, maybe there is a road map there?

I would guess they are"so good" because of the vast resources thrown at them, both in terms of compute and training data.

glc's avatar

With regard to pareidolia+theory of mind (as alluded to here): https://softwarecrisis.dev/letters/llmentalist/