Agentโ™ฅ๏ธŽAge
Catalog

FunASR

Official

by modelscope ยท Python

Transcribe local audio with FunASR and SenseVoice using private, on-device inference.

io.github.modelscope/funasr-mcp MCP Server

The Model Context Protocol (MCP) server provides local audio transcription using FunASR and SenseVoice, with SenseVoiceSmall as the default. It supports private, on-device inference and is positioned for speech-to-text workflows, including streaming ASR, multilingual ASR, punctuation, and speaker diarization.

๐Ÿ› ๏ธ Key Features

  • Local audio transcription via FunASR and SenseVoice
  • SenseVoiceSmall used by default
  • Private, on-device inference
  • Speech-to-text with options referenced in topics: streaming-asr, multilingual-asr, punctuation, speaker-diarization

๐Ÿš€ Use Cases

  • Transcribing audio on-device for Chinese and multilingual content
  • Building ASR applications that need real-time/streaming transcription
  • Generating structured transcripts with punctuation and diarization

โšก Developer Benefits

  • MCP server for directory checks and tool discovery (mentions tools/list)
  • Docker support: runs the MCP server over stdio
  • Runnable example setup with configurable device via FUNASR_DEVICE=cpu

โš ๏ธ Limitations

  • Source material describes setup and inference scope, but does not provide performance, accuracy, or audio constraints.

Topics

  • pytorch, speech-recognition, paraformer, punctuation, speaker-diarization, voice-activity-detection, asr, multilingual-asr, speech-to-text, transcription, whisper-alternative, audio, chinese, emotion-recognition, mcp-server, openai-compatible-api, streaming-asr, vllm, funasr, real-time-asr

Topics

pytorchspeech-recognitionparaformerpunctuationspeaker-diarizationvoice-activity-detectionasrmultilingual-asrspeech-to-texttranscriptionwhisper-alternativeaudiochineseemotion-recognitionmcp-serveropenai-compatible-apistreaming-asrvllmfunasrreal-time-asr