Metadata →
Highlights
id1030417322
To get real insight, you need to try using AI for different use cases and rigorously assess how good they are in the areas that matter to you.
A propósito de los benchmarks y evals de los LLM.
id1030434346
work is increasingly about assigning work to agents, rather than working together with chatbots. A joint study by OpenAI and academic economists shows how quickly this is happening inside their own organization. Critically, it isn’t just coders who are using agents. Legal, HR, and other non-tech functions have adopted agents at nearly the same rate. OpenAI may be a sort of canary in the coal mine for what will happen elsewhere in work.
id1030434449
What actually mattered was not the profession of the user, but their expertise. The more domain experience someone had, the more successful they were in using Claude Code in that domain. And, even more interestingly, the more useful output they got from Claude from each prompt.
Evidencia a favor de que los expertos en dominios específicos son los mejores posicionados para hacer un uso efectivo de los agentes de IA.
id1030434641
The instability is what happens when institutions that move at the speed of people (or worse, committees) try to track a capability curve that is very much not human in nature. And as long as we are on some sort of exponential, and for as long as it lasts, the gap only widens.
La exponencial → Situación actual en que las métricas que trackean la capacidad de los modelos va aumentando exponencialmente. Determina “la brecha” → la distancia entre las capacidades de los modelos y la de las instituciones (que se mueven a velocidad humana) para integrarlas de manera óptima en sus procesos.