Discussion about this post

User's avatar
Alex Tolley's avatar

If only lawyers would use Dan Davies' method to check that the genAI spitting out a brief to check the output. Maybe a few more expensive fines, perhaps even a disbarment, might change behaviors. In those sort of cases, I am not sure if a better way is to ask the AI to provide cases with references relevant to the case, with a summary of why each case law might be used. Then check everything, use the cases as narrow documents in a RAG, and ask the AI to create a brief only with those curated cases. Then read the brief carefully to make sure it is still correct.

GenAIs should start getting better at retrieving relevant references than the enshittified Google search that is now worse than it was. Surely Google understands that offering a free Gemini AI search summary is undermining their enshittified, ad-supported search results?

It may be somewhat recursive, but can genAI be asked to provide a good prompt for a desired output?

It seems to me that the best way to use genAI is like the way one uses a spreadsheet or code, use it to test out ideas quickly. A spreadsheet removes the slow, repetitive calculation task and quickly shows the result. GenAI should be able to quickly test out how to solve a task, which the user then curates/refines, to produce a good task output. Humans and AIs working together, with the AI acting like a fast go-fer assistant,

I don't know where AI goes next, other than to respond more effectively to prompts. I cannot help but feel that Asimov got there with controlling robots many decades ago when he wrote that the Spacers, with many robots to control, were far better at voice commands than the Earthers, who had little, if any, exposure to robots. We may just learn how best to interact with AIs to get the desired response. It would certainly be nice to use voice input to state whether accuracy or creativity was important in the response, perhaps using a human role as the initial prompt, just as we sort of use "You are a knowledgeable assistant on this subject [X]". I would prefer a catalog of domain specialist AIs, rather than trying to tailor the responses of a general AI. I also think that domain specialists may be able to sell their expertise as an off-the-shelf AI. A BradBot that can answer questions on economic history for a trifling amount might be very useful, with reviewers providing star ratings as feedback. I also would like to stage debates between domain expert AIs to determine which argument[s] are a better way forward, and to expose flawed logic. IOW, different AIs rather than a "mixture of experts" model.

George Kappus's avatar

Before I rely on an AI site for any task, I think it’s useful to try an experiment to see how it does on an obscure subject where you already know the answer to the query you’re going to post. An example: I spend a fair amount of time birding in Colombia. When I first tried Claude I posed it a query: What is the most critically endangered bird species in Colombia? I happened to know the answer to the question because I bid in Colombia with a scientist who is studying the species and working on its conservation. The answer is Antioquia Brushfinch, of which there are approximately 125 surviving adult individuals in the world. The identity of the brushfinch is not exactly a secret; there are lots of articles on its status. Yet Claude came back with an answer that just listed a set of widely known, somewhat endangered species. My follow up and their response went like this (paraphrasing : Hey, Bozo, what about Antioquia Brushfinch? Oh, sorry sir, we weren’t trying to answer the question you asked, just give some examples (all of which have populations over 125); after all who even really knows which species is most endangered. But, come to think about it, the brushfinch is pretty endangered and here are some more random examples. The moral of the story: Don’t trust AI unless they cite authority that you can confirm actually exists, is published in a reputable publication, actually says what they say it says and hasn’t been seriously called into question. In short, AI is only useful to the extent that it gets you quickly to some resource that does answer your question. (I’m assessing it only as a research to, not as a tool for writing code, for example.)

8 more comments...

No posts

Ready for more?