skills/modelscope.cn/Automatic Speech Recognition (ASR)

Automatic Speech Recognition (ASR)

Installation
SKILL.md

Automatic Speech Recognition (ASR)

Overview

After speaker diarization, you need to transcribe each speech segment to text. Whisper is the current state-of-the-art for ASR, with multiple model sizes offering different trade-offs between accuracy and speed.

When to Use

  • After speaker diarization is complete
  • Need to generate speaker-labeled transcripts
  • Creating subtitles from audio segments
  • Converting speech segments to text

Whisper Model Selection

Model Size Comparison

Installs
1
First Seen
Jun 7, 2026
Automatic Speech Recognition (ASR) from modelscope.cn