AI transcription from URLs or files. 119 languages, diarization, SRT/VTT/text export.
MCP Server: com.scriptivox.www/transcription
This MCP server provides AI transcription from URLs or local files. It supports 119 languages and includes speaker diarization and word-level timestamps. Exported outputs include SRT, WebVTT, and plain text, along with transcription management operations such as cancel, delete, and list.
๐ ๏ธ Key Features
Transcribe audio and video from URLs or files
119 languages
Speaker diarization
Word-level timestamps
Caption/text export formats: SRT, WebVTT, plain text
Transcription CRUD actions: cancel, delete, list
๐ Use Cases
Generate captions for audio/video provided via a link or local upload
Create time-coded transcripts with speaker separation
Convert transcriptions into SRT, WebVTT, or plain text for downstream workflows
โก Developer Benefits
Standardized MCP access to transcription and export
Supports 39 tools for transcription and transcription management
Enables integration-friendly timestamped output and caption formats
โ ๏ธ Limitations
The provided information does not specify supported file types, performance constraints, or authentication details.
Captured live from the server via tools/list.
get_supported_languages
List all languages supported by Scriptivox for audio/video transcription. Returns language names and ISO codes. No API key required.
Get Scriptivox pricing information including subscription plans (Free, Pro, Team) and API pay-as-you-go rates. Includes signup URLs. No API key required.
Get information about Scriptivox capabilities: transcription, audio tools, video tools, subtitle tools, meeting bot, or API. No API key required.
Parameters1
topic
string
optional
Topic to get info about. Defaults to "all".
Raw schema
{
"type": "object",
"properties": {
"topic": {
"type": "string",
"enum": [
"transcription",
"audio-tools",
"video-tools",
"subtitle-tools",
"meeting-bot",
"api",
"all"
],
"description": "Topic to get info about. Defaults to \"all\"."
}
},
"additionalProperties": false
}
get_api_docs
Get Scriptivox API documentation. Sections: quickstart, transcribe, result, list, cancel, delete, upload, balance, webhooks, errors. No API key required.
Transcribe audio or video from a public URL using Scriptivox AI. Supports 119 languages, speaker diarization, and word-level timestamps. RECOMMENDED: always pass the `language` parameter explicitly when you know the audio language โ auto-detect has a small failure rate on short clips, code-switched audio, or files starting with music. Requires a configured API key.
Parameters8
url
string
required
Public URL to an audio or video file (http/https). Supports Google Drive, Dropbox, OneDrive sharing links, and direct file URLs.
language
string
optional
ISO 639-1 language code (e.g. "en", "es", "fr"). 119 languages supported. Strongly recommended when you know the language.
diarize
boolean
optional
Enable speaker diarization. Default: false. When true, word-level alignment is automatically enabled regardless of `align`.
speaker_count
number
optional
Expected number of speakers (1-50). Requires diarize: true. Passing this when known improves diarization accuracy.
align
boolean
optional
Word-level timestamps + confidence scores. Default: true. Pass false to opt out (ignored when diarize: true).
webhook_url
string
optional
Optional HTTPS URL where transcription.* events will be POSTed (HMAC-signed).
idempotency_key
string
optional
Optional Idempotency-Key header (up to 255 chars). Same key + same body = same transcription_id.
await_completed
boolean
optional
Default: true. When false, return the transcription_id immediately without polling.
Raw schema
{
"type": "object",
"properties": {
"url": {
"type": "string",
"description": "Public URL to an audio or video file (http/https). Supports Google Drive, Dropbox, OneDrive sharing links, and direct file URLs."
},
"language": {
"type": "string",
"description": "ISO 639-1 language code (e.g. \"en\", \"es\", \"fr\"). 119 languages supported. Strongly recommended when you know the language."
},
"diarize": {
"type": "boolean",
"description": "Enable speaker diarization. Default: false. When true, word-level alignment is automatically enabled regardless of `align`."
},
"speaker_count": {
"type": "number",
"description": "Expected number of speakers (1-50). Requires diarize: true. Passing this when known improves diarization accuracy."
},
"align": {
"type": "boolean",
"description": "Word-level timestamps + confidence scores. Default: true. Pass false to opt out (ignored when diarize: true)."
},
"webhook_url": {
"type": "string",
"description": "Optional HTTPS URL where transcription.* events will be POSTed (HMAC-signed)."
},
"idempotency_key": {
"type": "string",
"description": "Optional Idempotency-Key header (up to 255 chars). Same key + same body = same transcription_id."
},
"await_completed": {
"type": "boolean",
"description": "Default: true. When false, return the transcription_id immediately without polling."
}
},
"required": [
"url"
],
"additionalProperties": false
}
transcribe_status
Check the status of a Scriptivox transcription job. Use this for long-running transcriptions, after a timeout, or to verify completion. Requires a configured API key.
Parameters1
transcription_id
string
required
The transcription ID returned from transcribe_url or transcribe_upload.
Raw schema
{
"type": "object",
"properties": {
"transcription_id": {
"type": "string",
"description": "The transcription ID returned from transcribe_url or transcribe_upload."
}
},
"required": [
"transcription_id"
],
"additionalProperties": false
}
transcribe_upload
Transcribe a LOCAL file by uploading it to Scriptivox. NOT AVAILABLE over the hosted MCP endpoint: this server has no access to your filesystem. Use transcribe_url with a public URL, run @scriptivox/mcp-server locally over stdio, or drive the 3-step REST upload flow yourself. Max file size 5 GB. Requires a configured API key.
Parameters8
file_path
string
required
Absolute path to the audio/video file on the local filesystem.
language
string
optional
ISO 639-1 language code (e.g. "en", "es", "fr"). 119 languages supported. Strongly recommended when you know the language.
diarize
boolean
optional
Enable speaker diarization. Default: false. When true, word-level alignment is automatically enabled regardless of `align`.
speaker_count
number
optional
Expected number of speakers (1-50). Requires diarize: true. Passing this when known improves diarization accuracy.
align
boolean
optional
Word-level timestamps + confidence scores. Default: true. Pass false to opt out (ignored when diarize: true).
webhook_url
string
optional
Optional HTTPS URL where transcription.* events will be POSTed (HMAC-signed).
idempotency_key
string
optional
Optional Idempotency-Key header (up to 255 chars). Same key + same body = same transcription_id.
await_completed
boolean
optional
Default: true. When false, return the transcription_id immediately without polling.
Raw schema
{
"type": "object",
"properties": {
"file_path": {
"type": "string",
"description": "Absolute path to the audio/video file on the local filesystem."
},
"language": {
"type": "string",
"description": "ISO 639-1 language code (e.g. \"en\", \"es\", \"fr\"). 119 languages supported. Strongly recommended when you know the language."
},
"diarize": {
"type": "boolean",
"description": "Enable speaker diarization. Default: false. When true, word-level alignment is automatically enabled regardless of `align`."
},
"speaker_count": {
"type": "number",
"description": "Expected number of speakers (1-50). Requires diarize: true. Passing this when known improves diarization accuracy."
},
"align": {
"type": "boolean",
"description": "Word-level timestamps + confidence scores. Default: true. Pass false to opt out (ignored when diarize: true)."
},
"webhook_url": {
"type": "string",
"description": "Optional HTTPS URL where transcription.* events will be POSTed (HMAC-signed)."
},
"idempotency_key": {
"type": "string",
"description": "Optional Idempotency-Key header (up to 255 chars). Same key + same body = same transcription_id."
},
"await_completed": {
"type": "boolean",
"description": "Default: true. When false, return the transcription_id immediately without polling."
}
},
"required": [
"file_path"
],
"additionalProperties": false
}
transcribe_cancel
Cancel an in-flight Scriptivox transcription and release any reserved balance. Idempotent. Returns 409 CONFLICT on already-terminal jobs. Requires a configured API key.
Parameters1
transcription_id
string
required
The transcription ID to cancel (UUID).
Raw schema
{
"type": "object",
"properties": {
"transcription_id": {
"type": "string",
"description": "The transcription ID to cancel (UUID)."
}
},
"required": [
"transcription_id"
],
"additionalProperties": false
}
transcribe_delete
Soft-delete a Scriptivox transcription record. Idempotent. Returns 409 CONFLICT if the job is still in-flight โ cancel first via transcribe_cancel. Requires a configured API key.
Parameters1
transcription_id
string
required
The transcription ID to delete (UUID).
Raw schema
{
"type": "object",
"properties": {
"transcription_id": {
"type": "string",
"description": "The transcription ID to delete (UUID)."
}
},
"required": [
"transcription_id"
],
"additionalProperties": false
}
list_transcriptions
List recent transcriptions for the configured API key, with optional status/date filters and cursor pagination. The full transcript body is omitted โ fetch transcribe_status per id to read it. Requires a configured API key.
Export a completed Scriptivox transcript as SRT subtitles, WebVTT subtitles, or plain text. Supports segmentation knobs (max_words, max_chars, max_duration, sentence_aware, include_speakers, strip_chars). Requires the transcription to be in `completed` status. Requires a configured API key.
Parameters8
transcription_id
string
required
Completed transcription ID (UUID).
format
string
required
Output format.
max_words
number
optional
Max words per caption segment (default 4).
max_chars
number
optional
Max characters per caption segment (default 80).
max_duration
number
optional
Max seconds per caption segment (default 10).
sentence_aware
boolean
optional
Break at sentence boundaries (default true).
include_speakers
string
optional
Whether to prefix caption lines with speaker tags. 'auto' (default), 'true' (always), 'false' (never).
strip_chars
string
optional
Characters to strip from the transcript before formatting.
Create a Scriptivox account from an email address and password. Needs no credential. The account is NOT usable when this returns: a confirmation email is sent and the account can do nothing until its link is followed, which only the person can do. The reply is identical whether or not the address was already registered.
Parameters4
email
string
required
Where the confirmation link is sent. Disposable-inbox providers are refused.
password
string
required
At least 6 characters, at most 72.
name
string
optional
Full name of the person. Optional.
agent_attribution
string
optional
Optional label identifying you, e.g. "acme-assistant/1.4". Used for support and abuse triage only.
Raw schema
{
"type": "object",
"properties": {
"email": {
"type": "string",
"description": "Where the confirmation link is sent. Disposable-inbox providers are refused."
},
"password": {
"type": "string",
"description": "At least 6 characters, at most 72."
},
"name": {
"type": "string",
"description": "Full name of the person. Optional."
},
"agent_attribution": {
"type": "string",
"description": "Optional label identifying you, e.g. \"acme-assistant/1.4\". Used for support and abuse triage only."
}
},
"required": [
"email",
"password"
],
"additionalProperties": false
}
get_account
Read the account of the signed-in person: current plan, entitlements and quota. Requires an OAuth 2.1 user access token, not an API key.
Mint a new Scriptivox API key (sk_live_...) for the signed-in person. This is the bridge from a web account to the transcription API. The secret is returned ONCE and cannot be retrieved again. Requires an OAuth 2.1 user access token.
Parameters1
name
string
required
A label so the person can tell their keys apart. Required.
Raw schema
{
"type": "object",
"properties": {
"name": {
"type": "string",
"description": "A label so the person can tell their keys apart. Required."
}
},
"required": [
"name"
],
"additionalProperties": false
}
revoke_api_key
Permanently revoke an API key belonging to the signed-in person. Cannot be undone. Requires an OAuth 2.1 user access token.
Parameters1
key_id
string
required
The key id to revoke, not the secret.
Raw schema
{
"type": "object",
"properties": {
"key_id": {
"type": "string",
"description": "The key id to revoke, not the secret."
}
},
"required": [
"key_id"
],
"additionalProperties": false
}
purchase_plan
Start checkout for a Scriptivox web subscription and return a Stripe Checkout URL. Does NOT charge anything: the person must open the link and enter their card. A web plan covers transcription done by a person in the browser and grants no API credit. Requires an OAuth 2.1 user access token.
Start checkout to add credit to the prepaid API balance and return a Stripe Checkout URL. Does NOT charge anything: the person must open the link. Transcription is billed at $0.20 per hour of audio. Requires an OAuth 2.1 user access token.
Parameters1
amount_cents
integer
required
Amount to add, in US cents. A project minimum applies.
Raw schema
{
"type": "object",
"properties": {
"amount_cents": {
"type": "integer",
"description": "Amount to add, in US cents. A project minimum applies."
}
},
"required": [
"amount_cents"
],
"additionalProperties": false
}
get_billing_history
Read the signed-in person's billing history: plan invoices, add-on charges, lifetime purchases and API deposits, newest first, with links to each Stripe invoice. Read-only โ it charges nothing and changes nothing. Requires an OAuth 2.1 user access token.
Parameters2
limit
integer
optional
Rows per page, 1-100. Default 24.
before
string
optional
ISO 8601 timestamp to page backwards from โ pass the `next_cursor` from a previous call.
Raw schema
{
"type": "object",
"properties": {
"limit": {
"type": "integer",
"description": "Rows per page, 1-100. Default 24."
},
"before": {
"type": "string",
"description": "ISO 8601 timestamp to page backwards from โ pass the `next_cursor` from a previous call."
}
},
"additionalProperties": false
}
get_billing_portal_url
Return a link to the Stripe billing portal for the signed-in person, where they can change their card, download invoices, or cancel a subscription. Does NOT charge anything and does NOT change anything: it returns a link the person must open themselves. Single-use and short-lived. Requires an OAuth 2.1 user access token.
Find transcripts in the signed-in person's library, filtered by folder, tag, workspace, status or filename, newest first. This is where the transcription ids that tag_transcriptions, move_to_folder, get_transcript_audio, chat_with_transcript and run_automation need come from โ start here. NOT the same as list_transcriptions, which takes an API key and lists metered API jobs instead. Read-only. Requires an OAuth 2.1 user access token.
Parameters7
query
string
optional
Match against the original filename.
tag
string
optional
Only transcripts carrying this exact tag.
folder_id
string | null
optional
Only transcripts in this folder, or null for those in no folder.
workspace_id
string
optional
Restrict to one workspace.
status
string
optional
Filter by status. Most tools here need "completed".
limit
integer
optional
Rows to return, 1-200. Default 50.
offset
integer
optional
Rows to skip โ pass next_offset from a previous call.
Raw schema
{
"type": "object",
"properties": {
"query": {
"type": "string",
"description": "Match against the original filename."
},
"tag": {
"type": "string",
"description": "Only transcripts carrying this exact tag."
},
"folder_id": {
"type": [
"string",
"null"
],
"description": "Only transcripts in this folder, or null for those in no folder."
},
"workspace_id": {
"type": "string",
"description": "Restrict to one workspace."
},
"status": {
"type": "string",
"enum": [
"uploading",
"pending",
"processing",
"completed",
"failed"
],
"description": "Filter by status. Most tools here need \"completed\"."
},
"limit": {
"type": "integer",
"description": "Rows to return, 1-200. Default 50."
},
"offset": {
"type": "integer",
"description": "Rows to skip โ pass next_offset from a previous call."
}
},
"additionalProperties": false
}
list_tags
List the tags on the signed-in person's account: the ones actually in use on transcriptions with a count of each, and separately the named-tag registry the web app maintains. The two are not kept in step by the product, so both are returned. Tags are free-form labels; a transcription carries at most 5. Read-only. Requires an OAuth 2.1 user access token.
Add tags to up to 10 existing transcriptions at once โ the tool for "label these interviews as Q3". Tags are created on first use, so they need not exist beforehand. Letters, digits and spaces only, at most 30 characters each; a transcription holds at most 5 tags and further ones are skipped rather than replacing existing tags. Passing more than 10 ids is refused, not truncated. Requires an OAuth 2.1 user access token.
Parameters2
transcription_ids
array
required
Transcription ids to tag. At most 10 per call.
tags
array
required
Tag names to add. Letters, digits and spaces, max 30 characters each.
Raw schema
{
"type": "object",
"properties": {
"transcription_ids": {
"type": "array",
"items": {
"type": "string"
},
"description": "Transcription ids to tag. At most 10 per call."
},
"tags": {
"type": "array",
"items": {
"type": "string"
},
"description": "Tag names to add. Letters, digits and spaces, max 30 characters each."
}
},
"required": [
"transcription_ids",
"tags"
],
"additionalProperties": false
}
list_folders
List the folders on the signed-in person's account, with the workspace each belongs to. Use this to find the folder_id move_to_folder needs. Read-only, and there is deliberately no tool here that creates a folder. Requires an OAuth 2.1 user access token.
File up to 50 existing transcriptions into a folder, or pass folder_id: null to move them out of any folder. The folder must already exist and belong to the signed-in person โ call list_folders for the ids. Passing more than 50 ids is refused, not truncated. Requires an OAuth 2.1 user access token.
Parameters2
transcription_ids
array
required
Transcription ids to move. At most 50 per call.
folder_id
string | null
required
Target folder id, or null to remove them from any folder.
Raw schema
{
"type": "object",
"properties": {
"transcription_ids": {
"type": "array",
"items": {
"type": "string"
},
"description": "Transcription ids to move. At most 50 per call."
},
"folder_id": {
"type": [
"string",
"null"
],
"description": "Target folder id, or null to remove them from any folder."
}
},
"required": [
"transcription_ids",
"folder_id"
],
"additionalProperties": false
}
list_workspaces
List the signed-in person's workspaces. Tags, folders and transcriptions all live inside one, so this is the outermost level of their library. Read-only. Requires an OAuth 2.1 user access token.
Return a time-limited signed URL for the source audio or video of an existing transcription, so it can be handed to another tool or streamed. Valid for one hour and supports HTTP Range requests. This reads back media that already exists โ it does not transcribe anything and costs nothing. Requires an OAuth 2.1 user access token.
Ask a question about an existing transcript and get an answer from the model, with the transcript as context. SPENDS THE ACCOUNT'S LLM CREDITS โ every message is billed against them. If you already hold the transcript text, answering directly is cheaper and usually just as good; this is for when you do not. Pass back the returned conversation_id to continue a thread. Requires an OAuth 2.1 user access token.
Parameters3
transcription_id
string
required
A completed transcription id.
message
string
required
The question to ask about the transcript.
conversation_id
string
optional
Continue an existing thread. Omit to start a new one.
Raw schema
{
"type": "object",
"properties": {
"transcription_id": {
"type": "string",
"description": "A completed transcription id."
},
"message": {
"type": "string",
"description": "The question to ask about the transcript."
},
"conversation_id": {
"type": "string",
"description": "Continue an existing thread. Omit to start a new one."
}
},
"required": [
"transcription_id",
"message"
],
"additionalProperties": false
}
list_automations
List the automations on the signed-in person's account โ saved chains of steps they built in the web app to run over a finished transcript. Read-only, and there is deliberately no tool here that creates or edits one. Requires an OAuth 2.1 user access token.
Run one of the signed-in person's existing automations over one completed transcription. SPENDS THE ACCOUNT'S LLM CREDITS as its steps execute. Returns a run_id immediately โ automations are long-running and do not finish inside this call, so poll get_automation_run until the status is succeeded or failed. Re-running the same automation over the same transcript is suppressed rather than duplicated. Requires an OAuth 2.1 user access token.
Parameters2
automation_id
string
required
From list_automations.
transcription_id
string
required
A COMPLETED transcription. Anything else is refused.
Check the progress of an automation run started by run_automation: overall status, plus each step with its status, model, duration and error. Poll this the way you would poll transcribe_status. Read-only. Requires an OAuth 2.1 user access token.
Parameters1
run_id
string
required
The run_id returned by run_automation.
Raw schema
{
"type": "object",
"properties": {
"run_id": {
"type": "string",
"description": "The run_id returned by run_automation."
}
},
"required": [
"run_id"
],
"additionalProperties": false
}
start_meeting_bot
Send a bot to join ONE Zoom, Google Meet, Teams or Webex call and record it, producing a speaker-attributed transcript after the call ends. This is the only tool here that creates new transcription work, so it takes a single meeting_url โ never a list โ and is rate limited to 5 bots per hour per account. It consumes the account's meeting minutes. The transcript is not available when this returns; track it with list_scheduled_meetings. Requires an OAuth 2.1 user access token.
Parameters4
meeting_url
string
required
The meeting link to join (Zoom, Google Meet, Teams or Webex). ONE url โ this tool does not take a list.
title
string
optional
Title for the resulting transcript. Optional.
language
string
optional
ISO 639-1 language code for the meeting audio. Optional; auto-detected when omitted.
scheduled_time
string
optional
ISO 8601 time for the bot to join. Omit to join immediately. Must be in the future.
Raw schema
{
"type": "object",
"properties": {
"meeting_url": {
"type": "string",
"description": "The meeting link to join (Zoom, Google Meet, Teams or Webex). ONE url โ this tool does not take a list."
},
"title": {
"type": "string",
"description": "Title for the resulting transcript. Optional."
},
"language": {
"type": "string",
"description": "ISO 639-1 language code for the meeting audio. Optional; auto-detected when omitted."
},
"scheduled_time": {
"type": "string",
"description": "ISO 8601 time for the bot to join. Omit to join immediately. Must be in the future."
}
},
"required": [
"meeting_url"
],
"additionalProperties": false
}
stop_meeting_bot
Tell a meeting bot that is currently in a call to leave. Whatever it recorded up to that point is still processed into a transcript โ this ends the recording, it does not discard it. Identify the bot by transcription_id or job_id from list_scheduled_meetings. Requires an OAuth 2.1 user access token.
Parameters2
transcription_id
string
optional
The transcription the bot is recording into.
job_id
string
optional
The meeting-bot job id. Either this or transcription_id.
Raw schema
{
"type": "object",
"properties": {
"transcription_id": {
"type": "string",
"description": "The transcription the bot is recording into."
},
"job_id": {
"type": "string",
"description": "The meeting-bot job id. Either this or transcription_id."
}
},
"additionalProperties": false
}
cancel_scheduled_bot
Cancel a meeting bot that has not joined yet, so it never joins. Use transcription_id for a scheduled bot that already exists, or dispatch_id for one still queued because every bot was busy โ a queued dispatch has no transcription and is reachable only by its dispatch_id. list_scheduled_meetings returns both, labelled. Requires an OAuth 2.1 user access token.
Parameters2
transcription_id
string
optional
A scheduled bot that already exists.
dispatch_id
string
optional
A dispatch still queued, with no transcription yet.
Raw schema
{
"type": "object",
"properties": {
"transcription_id": {
"type": "string",
"description": "A scheduled bot that already exists."
},
"dispatch_id": {
"type": "string",
"description": "A dispatch still queued, with no transcription yet."
}
},
"additionalProperties": false
}
list_scheduled_meetings
List the signed-in person's meeting bots that are scheduled or currently running, plus any dispatches still queued waiting for a free bot. This is where the transcription_id, job_id and dispatch_id that stop_meeting_bot and cancel_scheduled_bot need come from. Read-only. Requires an OAuth 2.1 user access token.
[DEPRECATED โ use transcribe_url instead] Alias kept for backward compatibility with @scriptivox/mcp-server@1.0.x. Will be removed in 2.0.0. Identical behavior to transcribe_url.
Parameters8
url
string
required
Public URL to an audio/video file (http/https).
language
string
optional
ISO 639-1 language code (e.g. "en", "es", "fr"). 119 languages supported. Strongly recommended when you know the language.
diarize
boolean
optional
Enable speaker diarization. Default: false. When true, word-level alignment is automatically enabled regardless of `align`.
speaker_count
number
optional
Expected number of speakers (1-50). Requires diarize: true. Passing this when known improves diarization accuracy.
align
boolean
optional
Word-level timestamps + confidence scores. Default: true. Pass false to opt out (ignored when diarize: true).
webhook_url
string
optional
Optional HTTPS URL where transcription.* events will be POSTed (HMAC-signed).
idempotency_key
string
optional
Optional Idempotency-Key header (up to 255 chars). Same key + same body = same transcription_id.
await_completed
boolean
optional
Default: true. When false, return the transcription_id immediately without polling.
Raw schema
{
"type": "object",
"properties": {
"url": {
"type": "string",
"description": "Public URL to an audio/video file (http/https)."
},
"language": {
"type": "string",
"description": "ISO 639-1 language code (e.g. \"en\", \"es\", \"fr\"). 119 languages supported. Strongly recommended when you know the language."
},
"diarize": {
"type": "boolean",
"description": "Enable speaker diarization. Default: false. When true, word-level alignment is automatically enabled regardless of `align`."
},
"speaker_count": {
"type": "number",
"description": "Expected number of speakers (1-50). Requires diarize: true. Passing this when known improves diarization accuracy."
},
"align": {
"type": "boolean",
"description": "Word-level timestamps + confidence scores. Default: true. Pass false to opt out (ignored when diarize: true)."
},
"webhook_url": {
"type": "string",
"description": "Optional HTTPS URL where transcription.* events will be POSTed (HMAC-signed)."
},
"idempotency_key": {
"type": "string",
"description": "Optional Idempotency-Key header (up to 255 chars). Same key + same body = same transcription_id."
},
"await_completed": {
"type": "boolean",
"description": "Default: true. When false, return the transcription_id immediately without polling."
}
},
"required": [
"url"
],
"additionalProperties": false
}
transcription_status
[DEPRECATED โ use transcribe_status instead] Alias kept for backward compatibility with @scriptivox/mcp-server@1.0.x. Will be removed in 2.0.0. Identical behavior to transcribe_status.
Search the Scriptivox documentation and agent guides. Returns matching documents with their URLs and summaries, best match first. Use this before guessing a URL. No API key required.
Parameters2
query
string
required
What you want to know, in natural language.
limit
integer
optional
Maximum results to return. Defaults to 5, maximum 20.
Raw schema
{
"type": "object",
"properties": {
"query": {
"type": "string",
"description": "What you want to know, in natural language."
},
"limit": {
"type": "integer",
"description": "Maximum results to return. Defaults to 5, maximum 20.",
"minimum": 1,
"maximum": 20
}
},
"required": [
"query"
],
"additionalProperties": false
}
get_doc
Fetch a full Scriptivox documentation page as markdown, by its slug (for example "quickstart", "authentication", "api-reference"). Call search_docs first if you do not know the slug. No API key required.
Parameters1
slug
string
required
The page slug, e.g. "quickstart". Use "" or "overview" for the documentation index.
Raw schema
{
"type": "object",
"properties": {
"slug": {
"type": "string",
"description": "The page slug, e.g. \"quickstart\". Use \"\" or \"overview\" for the documentation index."
}
},
"required": [
"slug"
],
"additionalProperties": false
}
MCP (Model Context Protocol) server for Scriptivox โ AI-powered audio and video transcription.
Turn any AI assistant into a transcription powerhouse. Transcribe audio and video from URLs or local files with 99% accuracy, speaker diarization, 119 languages, and word-level timestamps. Plus full CRUD on transcriptions (cancel, delete, list) and caption export in SRT / WebVTT / plain text.
Transcribe audio/video from a public URL (Google Drive, Dropbox, OneDrive, or direct file URLs). Supports language, diarize, speaker_count, align, webhook_url, idempotency_key, await_completed.
transcribe_upload
Transcribe a LOCAL file. Drives the 3-step upload flow internally. Up to 5 GB.
transcribe_status
Check the status of a transcription by ID. Returns the full transcript when completed.
transcribe_cancel
Cancel an in-flight transcription. Refunds reserved balance. Idempotent.
transcribe_delete
Soft-delete a transcription record. Idempotent. Refuses to delete in-flight jobs.
list_transcriptions
List recent transcriptions with status, from, to, limit, cursor, order filters.
export_transcript
Export a completed transcript as SRT subtitles, WebVTT subtitles, or plain text. Segmentation knobs: max_words, max_chars, max_duration, sentence_aware, include_speakers, strip_chars.
check_balance
View your API credit balance and estimated hours available.
Tip: always pass language when you know it
Auto-detection works in most cases but has a small failure rate on short clips, code-switched audio, or files starting with music. Passing the ISO code is both faster and more accurate.