6 Comments
User's avatar
StrangePolyhedrons's avatar

[You’re just not even going to waste your time asking an AI model to do something if you can only expect like a 50% success rate,” Witkin said.]

I guess it depends on how easy it is to check if it succeeded or failed. If I have to spend three hours asking it to keep trying until I get a winner, that's still better than 16 hours... but only if I can know if it's good or bad work pretty quickly.

Kobe Yank-Jacobs's avatar

Excellent point! I raise this in one of my questions

Willie Boag's avatar

Interesting point about the adoption of Codex at legal teams from the frontier labs! Nice way to disentangle the tool’s reliability vs the firm’s ability to integrate it into workflows

Aaron Bailey's avatar

Necro-comment, I know, but the piece I think you missed re: the “Canaries in the Coal Mine” paper is that entry level employees are already structurally assumed to be error prone and low-context. I’m in legal - any “job” (task, area of responsibility, etc.) that I’m giving to a 1st year associate or a junior paralegal is a job that either is already sufficiently insulated from real world consequences enough to insulate from AI mistakes, or that is already incorporated into workflows designed to error check against low-reliability junior staff.

Loren's avatar

This “conversation overheard in a bar” construction is rather awkward. In what kind of bar are there two academics poring through graphs casually over beer?

Neckberg's avatar

The relevant question isn't just "can we trust AI?", but "how trustworthy is AI compared to humans?".

And further, TRUST is basically just a nickname for PREDICTION, so the final question is "how PREDICTABLE is AI compared to humans?"

And probably, AI is actually more predictable than humans, or at least, we can likely create models that become more predictable in the long run.