Tools1.9k•06/28/2026, 12:15If you are using LLM-as-a-Judge for model evaluation, you should pay attention to this work.1 min readRead
GitHub1.6k•06/26/2026, 12:00Vesuvius Challenge Successfully Reads Charred Scroll for the First Time2 min readRead
Security200•08/30/2026, 11:00OpenAI Agents Created Three Civilizations Unknown to Humans4 min readRead