Tweet by zeeg

September 5, 2025

I think the fundamental issue is the evals are simply too unreliable and not scientific enough. Yes you could build crazy statistical test suite but iteration is a nightmare because don’t have a clue why shit fails. This is a real issue with LLM-based dev imo. https://t.co/aTezXJzY6S

Author
zeeg
Date
September 5, 2025