Gist
AI systems delivered consistent, low-cost ratings, but local clinicians still identified important concerns that automated evaluators overlooked. Study: Human evaluators vs. LLM-as-a-Judge: toward scalable evaluation of…