Agentβ™₯︎Age
Catalog

African Speech Corpora Quality Audit

Official

by papasega Β· Python

Read-only, provenance-aware selection of Wolof speech corpora by measured quality.

io.github.papasega/african-speech-corpora β€” African Speech Corpora MCP

A read-only, provenance-aware MCP server that provides a selection of Wolof speech corpora by measured quality. It implements the Model Context Protocol over public speech corpora for African languages, starting with Wolof, and also covering Pulaar and Sereer. The 2.0 catalog includes 14 variants.

πŸ› οΈ Key Features

  • Read-only MCP server
  • Provenance-aware selection of speech corpora
  • Measured quality–based selection
  • Wolof-first; Pulaar (ful) and Sereer (srr)
  • 2.0 catalog: 14 variants (11 original-source variants, 3 derivatives)

πŸš€ Use Cases

  • African-language speech dataset discovery
  • ASR-focused workflows using public corpora
  • Dataset audit and provenance-aware selection for Wolof-first needs
  • Support for low-resource languages beyond Wolof (Pulaar, Sereer)

⚑ Developer Benefits

  • Integrates with Model Context Protocol (MCP)
  • Structured access to dataset variants in a 2.0 catalog
  • Dataset-audit oriented selection via measured quality and provenance

⚠️ Limitations

  • Read-only access
  • Wolof-first catalog emphasis; Pulaar and Sereer support is secondary in the provided excerpt

Topics

african-languagesasraudio-datasetsdata-qualitydataset-audithuggingfacelow-resource-languagesmcpmcp-servermodel-context-protocolpythonspeech-datasetsspeech-recognitionttswolofwolof-nlp