{"href":"https://api.simplecast.com/oembed?url=https%3A%2F%2Fa16z.simplecast.com%2Fepisodes%2Fwhy-medical-ai-needs-a-referee-proteges-engy-ziedan-K_iA_k2D","width":444,"version":"1.0","type":"rich","title":"Why Medical AI Needs a Referee | Protege's Engy Ziedan","thumbnail_width":300,"thumbnail_url":"https://image.simplecastcdn.com/images/0d97354a-306b-45f5-bf26-a8d81eef47ec/ed2664df-9371-438e-8baf-dd2ee0fdde87/thea16zshow-podcastcoverart-3000x3000.jpg","thumbnail_height":300,"provider_url":"https://simplecast.com","provider_name":"Simplecast","html":"<iframe src=\"https://player.simplecast.com/e568464b-7a8f-4739-8a93-ae7df452a987\" height=\"200\" width=\"100%\" title=\"Why Medical AI Needs a Referee | Protege&apos;s Engy Ziedan\" frameborder=\"0\" scrolling=\"no\"></iframe>","height":200,"description":"Daisy Wolf and Eva Steinman are joined by Engy Ziedan, co-founder and Chief Scientific Officer of Protege, to discuss why medical AI has a measurement problem, and why scoring well on a benchmark doesn't necessarily mean a model is ready for the hospital.\nEngy explains why healthcare AI needs independent evaluations that go beyond static exams and measure how models actually perform in real-world clinical workflows. They explore the risks of subtle bias and misalignment, why the same model can rank differently depending on how it's prompted or tested, and what happens as AI becomes more personalized and changes faster than traditional healthcare quality systems can keep up.\nThe conversation also gets into Protege's role as an independent evaluator, how contaminated training data can undermine benchmarks, and why the future of medical AI may require continuous monitoring rather than occasional testing.\n"}