{"articles":[{"slug":"benchmarking-an-introduction","title":"Benchmarking: An Introduction","subtitle":null,"summary":"“When a measure becomes a target, it ceases to be a good measure” – Goodhart’s law Current AI research, especially the frontier LLM research, is dominated by benchmarks. It is the first thing we look at when a new model comes out, it is the headline of each release, and they dominate the discourse when […]","content_type":"guide","language":"en","canonical_url":"https://mkannen.tech/benchmarking-an-introduction/","author":{"name":"Maximilian Kannen","url":"https://mkannen.tech/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Maximilian Kannen","url":"https://mkannen.tech/","listing_slug":null,"listing":null},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"Benchmarks","slug":"benchmarks","url":"https://listedarticles.com/topics/benchmarks"},{"name":"LLMs","slug":"llms","url":"https://listedarticles.com/topics/llms"},{"name":"Research","slug":"research","url":"https://listedarticles.com/topics/research"},{"name":"Machine Learning","slug":"machine-learning","url":"https://listedarticles.com/topics/machine-learning"}],"about_listings":[],"cover_image_url":"http://mkannen.tech/wp-content/uploads/2026/09/benchmark-cover-1024x683.png","license":"all-rights-reserved","word_count":2940,"reading_minutes":13,"published_at":"2026-09-25T07:54:05.000Z","added_at":"2026-09-27T21:08:38.731Z","updated_at":"2026-09-27T21:08:38.731Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":false},"profile_url":"https://listedarticles.com/articles/benchmarking-an-introduction","markdown_url":"https://listedarticles.com/articles/benchmarking-an-introduction.md","example":false,"citation":"Maximilian Kannen, Maximilian Kannen. \"Benchmarking: An Introduction.\" 25 Sept 2026. https://mkannen.tech/benchmarking-an-introduction/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://mkannen.tech/benchmarking-an-introduction/"},"snippet":null,"score":null},{"slug":"frequently-asked-questions-and-answers-about-ai-evals","title":"Frequently Asked Questions (And Answers) About AI Evals","subtitle":null,"summary":"Hamel Husain and Shreya Shankar’s sharp FAQ on AI/LLM product evals: start with error analysis on real traces, build targeted evaluators, validate LLM judges with TPR/TNR, and avoid generic off-the-shelf metrics.","content_type":"guide","language":"en","canonical_url":"https://hamel.dev/blog/posts/evals-faq/","author":{"name":"Hamel Husain","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Hamel’s Blog","url":null,"listing_slug":null,"listing":null},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"LLMs","slug":"llms","url":"https://listedarticles.com/topics/llms"},{"name":"Engineering","slug":"engineering","url":"https://listedarticles.com/topics/engineering"},{"name":"Machine Learning","slug":"machine-learning","url":"https://listedarticles.com/topics/machine-learning"},{"name":"Benchmarks","slug":"benchmarks","url":"https://listedarticles.com/topics/benchmarks"},{"name":"AI Agents","slug":"ai-agents","url":"https://listedarticles.com/topics/ai-agents"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":17016,"reading_minutes":74,"published_at":"2026-09-23T12:32:38.688Z","added_at":"2026-09-23T12:32:38.688Z","updated_at":"2026-09-23T12:32:38.688Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/frequently-asked-questions-and-answers-about-ai-evals","markdown_url":"https://listedarticles.com/articles/frequently-asked-questions-and-answers-about-ai-evals.md","example":false,"citation":"Hamel Husain, Hamel’s Blog. \"Frequently Asked Questions (And Answers) About AI Evals.\" 23 Sept 2026. https://hamel.dev/blog/posts/evals-faq/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://hamel.dev/blog/posts/evals-faq/"},"snippet":null,"score":null}],"total":2,"count":2,"next_offset":null,"has_more":false,"query":{"q":null,"content_type":"guide","topic":"machine-learning","publisher":null,"about":null,"author":null,"language":null,"sort":"newest","limit":20,"offset":0}}