<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Speech Recognition (ASR) on Po-Wei Chen, MD — Physiatrist &amp; Builder</title><link>https://drpwchen.com/en/tags/speech-recognition-asr/</link><description>Recent content in Speech Recognition (ASR) on Po-Wei Chen, MD — Physiatrist &amp; Builder</description><image><title>Po-Wei Chen, MD — Physiatrist &amp; Builder</title><url>https://drpwchen.com/og-default.png</url><link>https://drpwchen.com/og-default.png</link></image><generator>Hugo</generator><language>en</language><lastBuildDate>Tue, 14 Jul 2026 21:40:00 +0800</lastBuildDate><atom:link href="https://drpwchen.com/en/tags/speech-recognition-asr/index.xml" rel="self" type="application/rss+xml"/><item><title>My data, my benchmark: Swapping the speech recognition (ASR) engine for my lecture note system</title><link>https://drpwchen.com/en/posts/my-data-my-benchmark/</link><pubDate>Mon, 13 Jul 2026 16:31:00 +0800</pubDate><guid>https://drpwchen.com/en/posts/my-data-my-benchmark/</guid><description>MediaTek&amp;#39;s official benchmark already proves Breeze-ASR-25 beats Whisper, but it isn&amp;#39;t testing your task. Using a 180,000-word glossary built from hundreds of textbooks as a yardstick, I found it captures 49% more real vocabulary and runs faster. Another finding not in any model card: the time alignment of the Taiwanese version, Breeze-ASR-26, has regressed. Default decoding only segments once every 30 seconds, which outright breaks when making subtitles or aligning with slides, but enabling word-level timestamps saves it. The testing method is open-sourced.</description></item></channel></rss>