1 minDan Shipper and Katie Parrott / Vibe Checkrss原文 ↗

Vibe Check: Claude Opus 5 Is Brilliant in Flashes, Frustrating in Practice

by Dan Shipper and Katie Parrott
in Vibe Check

Midjourney/Every illustration.
Was this newsletter forwarded to you? Sign up to get it in your inbox.

Claude Opus 5 had a strange first week at Every. It argued with instructions, stopped before the work was done, and fought the systems we’d built for earlier Claude models.
Then we deleted them.

With less process, Opus 5 sometimes got dramatically better. It built strong software, worked through bugs for hours, and produced more rigorous knowledge work. The less we told it how to work, the more capable it looked.

Our Vibe Check asks whether Opus 5 is disappointing—or whether the workflows it broke have become part of the problem. We tested it across coding, writing, knowledge work, and agents, with results that made the model hard to place.

Click here to read the full post
Want the full text of all articles in RSS? Become a subscriber, or learn more.

如果可以重来,你还愿意花这段时间读它吗?

· 匿名阅读记录只用于改进推荐