Studio
Open a media file from Projects, or load one locally.
AI Data & Model Operations Dashboard
Brimtek's workbench for multilingual speech, language, vision, model evaluation, call-center AI, media localization, conversational AI, and production delivery. The project and media selectors control the active operational context.
Core business
Speech & Voice AI
ASR evaluation, diarization, telephony transcription, dialect and domain testing, TTS evaluation, human gold data, and custom speech model workflows.
Media AI & Localization
Professional transcription, subtitles, dubbing, translation, live captions, media indexing, and human review for multilingual content.
Conversational AI
Call-center AI, voicebots, chatbots, multilingual conversation data, response evaluation, safety review, and domain adaptation.
Vision & Multimodal AI
Image and video labeling, scene and object annotation, OCR, video indexing, CV or VLM evaluation, and cross-modal quality review.
Operational coverage
Current project / media
Video Indexing Lab
Transcript / dub snapshot
Workbench entry points
Load the 15-clip Swedish speech-labeling demo set to inspect annotation, metadata, QA and export workflows.
Open Speech / Audio labeling for manual segmentation, transcription correction, tags and speaker metadata.
Re-open a native Brimtek annotation session for continued labeling or evaluation.
Domain / topic coverage (this session)
Platform capabilities
- Speech data & ASR operations - OpenAI diarized transcription, high-accuracy gpt-transcribe hybrid, remote WhisperX/pyannote, and imported diarization all land in one Brimtek schema.
- Transcript normalization & adjudication - cached punctuation, casing, Arabic orthography and proper-name cleanup with glossary/context; raw ASR is preserved.
- Versioned annotation sets - ASR, subtitle and dub groupings coexist; re-segmenting no longer destroys the previous cut.
- Language data & translation - separate subtitle translation and spoken dub adaptation are generated with shared context and cached.
- TTS / voice evaluation & synthesis - OpenAI expressive TTS/custom voice, Azure Speech, and remote Brimtek XTTS; delivery instructions can specify accent, tone, emotion and pacing.
- Realtime speech workflows - browser audio streams through Brimtek's server-side WebSocket proxy so the OpenAI key never reaches the browser. Public capture requires HTTPS.
- Computer vision / video intelligence - scene-change + periodic sampling, cached transcript context and OpenAI vision produce timestamped shot/entity/topic/on-screen-text records.
Local VAD vs. server transcription
The built-in ✨ Auto-segment in Label Audio is still a local WebAudio energy-VAD. It is useful for manual labeling because it detects speech/silence without uploading anything, but it does not transcribe or identify speakers.
For production media, use Studio → ① Transcribe. The server can now run OpenAI diarized ASR, the high-accuracy gpt-transcribe hybrid, or your remote WhisperX/pyannote service, then cache and context-refine the result before deriving subtitle or dub segment sets.
Speech / Audio Labeling
Load a clip, click ✨ Auto-segment to detect speech regions automatically, then click into any transcript box to type. Speech over 15s auto-splits.
Shortcuts: space play · S split · M merge · A auto-segment · J/K next/prev · Del remove
Source audio
Segments (0)
ASR Import / Advanced
Normal server transcription now lives in Studio and follows the active project/media automatically. Use this page only for an external GPU endpoint, Whisper/WhisperX JSON, SRT/VTT import, or other advanced/manual ingestion.
Transcription endpoint
not configured{audio, model, language} and expects Whisper-style {segments:[{start,end,text,speaker?}]} back.Or import an existing transcript
Pipeline - one raw transcript, many outputs
Transcribe (here)
Audio → time-stamped transcript via your ASR endpoint, or import existing Whisper/SRT output.
Post-edit & tag → Label Audio
Transcribers correct the machine output and add non-speech/markup tags. The 50-70% margin lever: verify, don't transcribe from scratch.
Translate → Translate / TTS QA
Per-segment translation with length-bounds validation for TTS fit.
Subtitle → Export SRT/VTT
Caption-ready export with the CPS/duration readability rule applied.
Dub → Dubbing / Side-by-Side
Source + target text side by side per segment, with per-row playback and synth hooks.
Live CC / Live Translation
Stream microphone or browser/tab audio through the Brimtek server. The OpenAI key stays on the droplet. The active project automatically supplies source/target language and context.
Source / live captions
Translated captions
Notes
Realtime transcription is optimized for low-latency captions and does not provide speaker labels. For final delivery, run the recorded media through the diarized file-transcription pipeline afterward. For browser/tab capture your browser may require selecting “share audio”.
Dubbing / Side-by-Side
Source ↔ target per segment, with per-row playback, merge (same-speaker guarded), split-at-second, and synth hooks. Load the media in Label Audio (or here), and rows from a CSV or the Raw Processing tab.
Segments (0)
Translate / TTS QA
Per-segment translation review, mirroring the source → target workflow with a TTS "Now Playing" strip for reviewing synthesized voice output against the transcript.
Segments (0)
No speech segments yet - add some in Label Audio first, then come back here.
Label Images
Computer-vision labeling workspace: draw bounding boxes, assign classes and attributes, and export training/evaluation annotations to COCO JSON or YOLO txt.
Annotations (0)
No boxes yet - draw on the image above.
Label Video
Video / vision-model labeling workspace: temporal actions, frame-level gesture regions, object/track metadata and validation-ready annotations for training or model evaluation.
Video annotations (0)
No annotations yet.
Subtitle & Dub Studio
Applied media-delivery workspace for documentary subtitling and dubbing. Works on the segments from Label Audio - QC against broadcast readability rules, build the dub script with duration budgets, export to delivery formats.
Dub pipeline - Colab interop (parameters mirror your notebooks)
🎙 Voice Studio - TTS & cloning
💰 Job pricing - cost vs. Azure list
Cue list & dub script (0)
No segments yet - create them in Label Audio (or import an SRT above).
Author Prompts
Write prompt-completion pairs to the Brimtek SFT schema - the 2023-style production workspace. RTL-aware, live word/char counts, per-item validation badge.
No items yet - click + New item, or import a .jsonl of {"prompt","completion",...} objects.
Evaluate Models
Explore a broad global model catalog, build a benchmark shortlist, then run human rubric scoring and pairwise preference judgments across language-model, speech, vision, and multimodal outputs.
Explore model catalog Hugging Face Hub first layer
Benchmark shortlist
0 selectedExport ratings
QA & Agreement
Inter-annotator agreement and gold-set scoring - the numbers that go in every client deliverable. Compare two annotators' work, or score an annotator against a gold standard.
Transcription agreement (two session JSONs)
Eval rating agreement (two ratings JSONLs from the Evaluate Models tab)
Reports: per-dimension exact agreement %, mean absolute difference (1-5 scale), and Cohen's κ for pairwise A/B/Tie choices. These are the statistics quoted in the Eval Pilot deliverable.
Validation Rules
Production QA gate - run all rules before delivery. Rules with no applicable data show N/A rather than failing.
Rule thresholds
Constraint settings (transcription conventions)
File Metadata & Speaker Roster
Required file-level values and per-speaker labels, per section 4 of the transcription guidelines.
File-level values
Speaker roster
Tag Reference
Full non-speech, markup, and disfluency tag set consolidated from the en_US v3.0 transcription guidelines and the IFAP 16kHz addendum.
Non-speech noise tags - human vocal
| [breath] | Inhalation/exhalation between words, yawning |
| [cough] | Coughing, throat clearing, sneezing |
| [cry] | Crying / sobbing |
| [laugh] | Laughing, chuckling |
| [lipsmack] | Lipsmacks, tongue-clicks |
Non-speech / non-human noise tags
| [applause] | Clapping - placed exactly where it occurred (16kHz addendum) |
| [beep] | Beep replacing profanity or classified info |
| [bg-speech] | Background speech overlapping the foreground speaker |
| [click] | Machine or phone click |
| [dtmf] | Telephone keypad tone |
| [music] | ≥1s of music/singing with no foreground speech; on-hold music |
| [no-speech] | Pause / silence of ≥1s, even with faint foreground noise mixed in |
| [noise] | Any other miscellaneous noise (screaming, rain, punching, etc.) |
| [ring] | Telephone ring |
| [sta] | Continuous static |
Markup tags
| <initial></initial> | Wraps an initialism/acronym spoken letter-by-letter, e.g. <initial>TV</initial> |
| <lang:Foreign></lang:Foreign> | Wraps a word/phrase in an unidentified non-target language |
| <lang:X></lang:X> | X = any capitalized language name (Arabic, Korean, Spanish, English…) |
| <overlap></overlap> | Wraps speech that overlaps with another speaker |
Disfluency & other markup
| #eh / #mm / #öh | Filler words - prefixed with # |
| word~ | Stumbled / truncated word, trailing tilde, no space before it |
| (( )) | Unintelligible speech - double parentheses, no inner spaces |
| ((word)) | Best-guess / low-confidence transcription of a barely intelligible word |
Rule: never insert a non-speech tag mid-word - place it immediately before the word in progress. Repeated instances of the same sound in one segment are represented only once.
Segmentation rules (quick reference)
- Five primary segment types: Speech, Babble, Overlap, Music, Noise. Each segment has exactly one primary type.
- Timestamps are positive floats in seconds.milliseconds (e.g. 12.345).
- Trim continuous silence/white noise ≥2s from the edges of a segment - keep segments tight around the target sound.
- Speech segments should not exceed 15 seconds - this workbench enforces it by auto-splitting on commit; use Split/Merge to fine-tune boundaries.
- Transcription is required only for Speech (and Babble/Overlap where applicable) segments.
- Single-channel files: segments with no speech or only non-speech sound on that channel should be dropped, not transcribed as [no-speech].
Export
Export the current session in the format your pipeline needs. Everything runs locally in the browser.
QA summary before export
Roadmap - one console, every modality
The metadata model (file-level values + annotator + typed segments) and the export pipeline are designed to extend beyond audio without a rewrite.
🖼 Photo / image labeling
Next upType.name: "IMAGE_ANNOTATION". Replaces time-based Segments with region-based Annotations.
- Bounding boxes, polygons, and segmentation masks
- Classification tags & multi-label scene/object tags
- Attribute tags (occlusion, truncation, blur, low-light)
- EXIF / capture metadata
🎬 Video labeling
PlannedType.name: "VIDEO_ANNOTATION". Segments gain a spatial component per frame range.
- Temporal action segments - same StartTime/EndTime pattern
- Object tracking IDs (reuses SpeakerId pattern as TrackId)
- Shot/scene boundary tags, camera motion tags
💬 Chatbot prompt / conversation labeling
PlannedType.name: "CONVERSATION_ANNOTATION". Segments become conversation turns.
- Turn-level role tags, intent & topic tags
- Response quality rubric scoring, A/B preference ranking
- Safety/policy category tags, red-team flags
🧩 Shared platform layer
Foundation - live todayWhat every future tab shares with the audio workbench built today:
- Guidelines/AddendumName versioning per file
- AnnotatorId + QA warning pipeline
- Consistent native-JSON envelope
- Same multi-format export layer
Projects
Each client job is a project. Upload media, run transcription on the server, then edit in Label Audio.
New project
All projects
Sign in to load projects.
Add media - select a project below first
Jobs
Server-side transcription. Auto-refreshes every 3 seconds while this tab is open.
Sign in to load jobs.
Voices & Consent
Every cloned voice must be linked to a signed release. Reference-audio upload is blocked server-side without recorded consent.
Register a voice
Registry
Sign in to load the registry.
Users
Admin only. Roles: admin, reviewer, annotator, client.
Add user
All users and credits
Service credit rates Internal Brimtek units, configurable by admin
Loading rates...