Brimtek Labs no media loaded File Status:Draft
offline Credits

Studio

Open a media file from Projects, or load one locally.

0.000 / 0.000
PIPELINE
0 segments

AI Data & Model Operations Dashboard

Brimtek's workbench for multilingual speech, language, vision, model evaluation, call-center AI, media localization, conversational AI, and production delivery. The project and media selectors control the active operational context.

Core business

01 · SPEECH

Speech & Voice AI

ASR evaluation, diarization, telephony transcription, dialect and domain testing, TTS evaluation, human gold data, and custom speech model workflows.

02 · MEDIA

Media AI & Localization

Professional transcription, subtitles, dubbing, translation, live captions, media indexing, and human review for multilingual content.

03 · CONVERSATION

Conversational AI

Call-center AI, voicebots, chatbots, multilingual conversation data, response evaluation, safety review, and domain adaptation.

04 · VISION

Vision & Multimodal AI

Image and video labeling, scene and object annotation, OCR, video indexing, CV or VLM evaluation, and cross-modal quality review.

Operational coverage

Multilingualhuman data and evaluation
Regional dialectsTelephony / call-center16 kHz mediaBroadcast and captionsASR evaluationTTS evaluationChatbot / voicebot QAVision / VLM labeling
Brimtek works with a mix of client-controlled, Brimtek-created, licensed, and permitted public-source material. Rights and permitted uses are tracked by project and contract.

Current project / media

Choose a project from the top bar.

Video Indexing Lab

Scene cuts, extracted frames, transcript context, vision analysis, OCR, entities and searchable timecodes.
No active indexed video.

Transcript / dub snapshot

No active transcript.
0
Segments labeled
0.0s
Total duration
0
Speakers tagged
0
QA warnings
-
Audio-min labeled / work-hour
-
Segments / hour
0
Eval items rated
-
Eval items / hour

Workbench entry points

Load the 15-clip Swedish speech-labeling demo set to inspect annotation, metadata, QA and export workflows.

Open Speech / Audio labeling for manual segmentation, transcription correction, tags and speaker metadata.

Re-open a native Brimtek annotation session for continued labeling or evaluation.

Domain / topic coverage (this session)

No segments yet.

Platform capabilities

  • Speech data & ASR operations - OpenAI diarized transcription, high-accuracy gpt-transcribe hybrid, remote WhisperX/pyannote, and imported diarization all land in one Brimtek schema.
  • Transcript normalization & adjudication - cached punctuation, casing, Arabic orthography and proper-name cleanup with glossary/context; raw ASR is preserved.
  • Versioned annotation sets - ASR, subtitle and dub groupings coexist; re-segmenting no longer destroys the previous cut.
  • Language data & translation - separate subtitle translation and spoken dub adaptation are generated with shared context and cached.
  • TTS / voice evaluation & synthesis - OpenAI expressive TTS/custom voice, Azure Speech, and remote Brimtek XTTS; delivery instructions can specify accent, tone, emotion and pacing.
  • Realtime speech workflows - browser audio streams through Brimtek's server-side WebSocket proxy so the OpenAI key never reaches the browser. Public capture requires HTTPS.
  • Computer vision / video intelligence - scene-change + periodic sampling, cached transcript context and OpenAI vision produce timestamped shot/entity/topic/on-screen-text records.

Local VAD vs. server transcription

The built-in ✨ Auto-segment in Label Audio is still a local WebAudio energy-VAD. It is useful for manual labeling because it detects speech/silence without uploading anything, but it does not transcribe or identify speakers.

For production media, use Studio → ① Transcribe. The server can now run OpenAI diarized ASR, the high-accuracy gpt-transcribe hybrid, or your remote WhisperX/pyannote service, then cache and context-refine the result before deriving subtitle or dub segment sets.

Speech / Audio Labeling

Load a clip, click ✨ Auto-segment to detect speech regions automatically, then click into any transcript box to type. Speech over 15s auto-splits.
Shortcuts: space play · S split · M merge · A auto-segment · J/K next/prev · Del remove

Source audio

Drop a .wav/.mp3 file here, or click to browse

Segments (0)

Selected segment

Click a segment row to edit its Start/End, Speaker, and Type here.

ASR Import / Advanced

Normal server transcription now lives in Studio and follows the active project/media automatically. Use this page only for an external GPU endpoint, Whisper/WhisperX JSON, SRT/VTT import, or other advanced/manual ingestion.

Transcription endpoint

not configured
Uses the audio loaded in Label Audio.
Security: a token in a browser file is visible to anyone who opens the file. Fine for your own machine; for shared/client use, point this at a small proxy on your droplet that holds the token server-side. The button POSTs {audio, model, language} and expects Whisper-style {segments:[{start,end,text,speaker?}]} back.

Or import an existing transcript

Accepts {segments:[…]} with start/end/text and optional speaker.
Timecodes + text parsed into segments.
No transcript loaded yet.

Pipeline - one raw transcript, many outputs

1

Transcribe (here)

Audio → time-stamped transcript via your ASR endpoint, or import existing Whisper/SRT output.

2

Post-edit & tag → Label Audio

Transcribers correct the machine output and add non-speech/markup tags. The 50-70% margin lever: verify, don't transcribe from scratch.

3

Translate → Translate / TTS QA

Per-segment translation with length-bounds validation for TTS fit.

4

Subtitle → Export SRT/VTT

Caption-ready export with the CPS/duration readability rule applied.

5

Dub → Dubbing / Side-by-Side

Source + target text side by side per segment, with per-row playback and synth hooks.

Live CC / Live Translation

Stream microphone or browser/tab audio through the Brimtek server. The OpenAI key stays on the droplet. The active project automatically supplies source/target language and context.

idle

Source / live captions

Translated captions

Notes

Realtime transcription is optimized for low-latency captions and does not provide speaker labels. For final delivery, run the recorded media through the diarized file-transcription pipeline afterward. For browser/tab capture your browser may require selecting “share audio”.

Dubbing / Side-by-Side

Source ↔ target per segment, with per-row playback, merge (same-speaker guarded), split-at-second, and synth hooks. Load the media in Label Audio (or here), and rows from a CSV or the Raw Processing tab.

Segments (0)

Translate / TTS QA

Per-segment translation review, mirroring the source → target workflow with a TTS "Now Playing" strip for reviewing synthesized voice output against the transcript.

Now Playing - no clip loaded

Segments (0)

No speech segments yet - add some in Label Audio first, then come back here.

Label Images

Computer-vision labeling workspace: draw bounding boxes, assign classes and attributes, and export training/evaluation annotations to COCO JSON or YOLO txt.

Drop a .jpg/.png here, or click to browse

Annotations (0)

No boxes yet - draw on the image above.

Label Video

Video / vision-model labeling workspace: temporal actions, frame-level gesture regions, object/track metadata and validation-ready annotations for training or model evaluation.

Drop a .mp4/.webm here, or click to browse

Video annotations (0)

No annotations yet.

Subtitle & Dub Studio

Applied media-delivery workspace for documentary subtitling and dubbing. Works on the segments from Label Audio - QC against broadcast readability rules, build the dub script with duration budgets, export to delivery formats.

Run QC to check every cue against the preset.

Dub pipeline - Colab interop (parameters mirror your notebooks)

Import a Whisper/diarization JSON, QC it here, then export the XTTS CSV your notebook already reads.

🎙 Voice Studio - TTS & cloning

Auto-fit computes the speaking-rate adjustment each cue needs to land inside its time slot, and writes it into the SSML - the core problem in dubbing. Cloning requires a signed release for every voice (Contract Pack, Doc A/C).

💰 Job pricing - cost vs. Azure list

Enter job size and press Calculate.
Azure figures are placeholders - their pricing page loads values dynamically. Put your region's real numbers from the Azure pricing calculator in these fields before quoting.

Cue list & dub script (0)

No segments yet - create them in Label Audio (or import an SRT above).

Author Prompts

Write prompt-completion pairs to the Brimtek SFT schema - the 2023-style production workspace. RTL-aware, live word/char counts, per-item validation badge.

0 / 0 0 errors

No items yet - click + New item, or import a .jsonl of {"prompt","completion",...} objects.

Evaluate Models

Explore a broad global model catalog, build a benchmark shortlist, then run human rubric scoring and pairwise preference judgments across language-model, speech, vision, and multimodal outputs.

Explore model catalog Hugging Face Hub first layer

Checking Hugging Face connection...Catalog browsing does not run inference or spend Brimtek credits. Adding a model creates a benchmark shortlist only.
Search the Hugging Face Hub to discover public, gated, private-access, and provider-served models available to your configured account.

Benchmark shortlist

0 selected
No models selected yet.
Rubric item: {"id","prompt","response","language"} · Pairwise item: {"id","prompt","response_a","response_b","language"}
3 rubric + 3 pairwise, multilingual - for training raters and demoing to clients.

Export ratings

No items loaded.

QA & Agreement

Inter-annotator agreement and gold-set scoring - the numbers that go in every client deliverable. Compare two annotators' work, or score an annotator against a gold standard.

Transcription agreement (two session JSONs)

Eval rating agreement (two ratings JSONLs from the Evaluate Models tab)

Reports: per-dimension exact agreement %, mean absolute difference (1-5 scale), and Cohen's κ for pairwise A/B/Tie choices. These are the statistics quoted in the Eval Pilot deliverable.

Validation Rules

Production QA gate - run all rules before delivery. Rules with no applicable data show N/A rather than failing.

-
Pass
-
Warn
-
Fail
-
N/A

Rule thresholds

Constraint settings (transcription conventions)

File Metadata & Speaker Roster

Required file-level values and per-speaker labels, per section 4 of the transcription guidelines.

File-level values

Speaker roster

Tag Reference

Full non-speech, markup, and disfluency tag set consolidated from the en_US v3.0 transcription guidelines and the IFAP 16kHz addendum.

Non-speech noise tags - human vocal

[breath]Inhalation/exhalation between words, yawning
[cough]Coughing, throat clearing, sneezing
[cry]Crying / sobbing
[laugh]Laughing, chuckling
[lipsmack]Lipsmacks, tongue-clicks

Non-speech / non-human noise tags

[applause]Clapping - placed exactly where it occurred (16kHz addendum)
[beep]Beep replacing profanity or classified info
[bg-speech]Background speech overlapping the foreground speaker
[click]Machine or phone click
[dtmf]Telephone keypad tone
[music]≥1s of music/singing with no foreground speech; on-hold music
[no-speech]Pause / silence of ≥1s, even with faint foreground noise mixed in
[noise]Any other miscellaneous noise (screaming, rain, punching, etc.)
[ring]Telephone ring
[sta]Continuous static

Markup tags

<initial></initial>Wraps an initialism/acronym spoken letter-by-letter, e.g. <initial>TV</initial>
<lang:Foreign></lang:Foreign>Wraps a word/phrase in an unidentified non-target language
<lang:X></lang:X>X = any capitalized language name (Arabic, Korean, Spanish, English…)
<overlap></overlap>Wraps speech that overlaps with another speaker

Disfluency & other markup

#eh / #mm / #öhFiller words - prefixed with #
word~Stumbled / truncated word, trailing tilde, no space before it
(( ))Unintelligible speech - double parentheses, no inner spaces
((word))Best-guess / low-confidence transcription of a barely intelligible word

Rule: never insert a non-speech tag mid-word - place it immediately before the word in progress. Repeated instances of the same sound in one segment are represented only once.

Segmentation rules (quick reference)

  • Five primary segment types: Speech, Babble, Overlap, Music, Noise. Each segment has exactly one primary type.
  • Timestamps are positive floats in seconds.milliseconds (e.g. 12.345).
  • Trim continuous silence/white noise ≥2s from the edges of a segment - keep segments tight around the target sound.
  • Speech segments should not exceed 15 seconds - this workbench enforces it by auto-splitting on commit; use Split/Merge to fine-tune boundaries.
  • Transcription is required only for Speech (and Babble/Overlap where applicable) segments.
  • Single-channel files: segments with no speech or only non-speech sound on that channel should be dropped, not transcribed as [no-speech].

Export

Export the current session in the format your pipeline needs. Everything runs locally in the browser.

Transcription
Native JSON
Full schema - Type, MetaData, Speakers, Segments with TranscriptionData. Matches your delivery format. Re-importable from the Dashboard.
HF / Whisper manifest (.jsonl)
One JSON object per line: audio_filepath, text (ASR-ready), duration, language, speaker.
CSV
segment_id, start, end, duration, speaker, language, type, transcription_raw, transcription_clean.
SRT subtitles
Speech segments as SRT captions (ASR-ready text) for sync with video review tools.
WebVTT
Same as SRT but WebVTT, for HTML5 <track> / web players.
Praat TextGrid
A single IntervalTier ("segments") with all boundaries and clean text - opens directly in Praat.
Audacity label track
Tab-separated start\tend\tlabel - import via Audacity's File → Import → Labels.
Kaldi-style bundle
wav.scp, segments, text, utt2spk, spk2gender - zipped. Loads JSZip once; falls back to loose files offline.
Translation
Bilingual CSV
segment_id, start, end, source_text, target_language, translation.
Translation JSON
Same native-JSON envelope, with a Translation object added to each segment.
Full project bundle
Project folder (.zip)
/audio, /transcripts (json+csv+srt+vtt), /translations, /metadata, /exports (kaldi+textgrid+audacity) - one zip, ready to hand off.

QA summary before export

No segments yet.

Roadmap - one console, every modality

The metadata model (file-level values + annotator + typed segments) and the export pipeline are designed to extend beyond audio without a rewrite.

🎬 Video labeling

Planned

Type.name: "VIDEO_ANNOTATION". Segments gain a spatial component per frame range.

  • Temporal action segments - same StartTime/EndTime pattern
  • Object tracking IDs (reuses SpeakerId pattern as TrackId)
  • Shot/scene boundary tags, camera motion tags

💬 Chatbot prompt / conversation labeling

Planned

Type.name: "CONVERSATION_ANNOTATION". Segments become conversation turns.

  • Turn-level role tags, intent & topic tags
  • Response quality rubric scoring, A/B preference ranking
  • Safety/policy category tags, red-team flags

🧩 Shared platform layer

Foundation - live today

What every future tab shares with the audio workbench built today:

  • Guidelines/AddendumName versioning per file
  • AnnotatorId + QA warning pipeline
  • Consistent native-JSON envelope
  • Same multi-format export layer

Projects

Each client job is a project. Upload media, run transcription on the server, then edit in Label Audio.

New project

All projects

Sign in to load projects.

Add media - select a project below first

You are responsible for holding the rights to any media you ingest (Work Order clause 5).
…or drop a video/audio file here, or click to browse

Jobs

Server-side transcription. Auto-refreshes every 3 seconds while this tab is open.

Sign in to load jobs.

Voices & Consent

Every cloned voice must be linked to a signed release. Reference-audio upload is blocked server-side without recorded consent.

Register a voice

Registry

Sign in to load the registry.

Users

Admin only. Roles: admin, reviewer, annotator, client.

Add user

Existing users remain unmetered until an admin enables enforcement.

All users and credits

Service credit rates Internal Brimtek units, configurable by admin

Loading rates...

Change my password

BRIMTEK
Human intelligence
for a brighter AI world
Trusted data. Better AI.

Human Data & Model Evaluation for Multilingual AI

Brimtek helps organizations build, evaluate, and improve speech, language, vision, and multimodal systems. Our workbench supports call-center AI, chatbots, subtitles, dubbing, live captions, security workflows, and production model evaluation.

Speech & Voice AITranscription, speaker labeling, diarization, telephony, TTS evaluation, dialect testing, and domain adaptation.
Media AI & LocalizationSubtitles, dubbing, translation, live captions, media indexing, and human review.
Conversational AICall-center AI, voicebots, chatbots, multilingual evaluation, dialogue data, and safety review.
Vision & Multimodal AIImage and video labeling, scene understanding, OCR, indexing, and CV or VLM evaluation.
Already working with Brimtek?Access projects, labeling queues, model evaluations, credits, media workflows, and exports from the secure Workbench.
VIDEO INDEX DEMO Real image frames · scene cuts · OCR · entities · transcript alignment FRAME 01 / 05
Extract frames
Detect scenes
Vision + OCR
Align transcript
Build index
Public Cambodia indexing demo frame
00:00.000 · establishing shot
A sunrise establishes place and time before the narration begins.
0:00 / 0:39
A   Source
Made in Cambodia is a story of people, craft and opportunity.
00:00.00 - 00:05.40
B   Arabic
صنع في كمبوديا هو قصة الناس والحرف والفرص.
aligned
C   Indonesian
Dibuat di Kamboja adalah kisah tentang manusia, kerajinan, dan peluang.
aligned
Public demo sequence with real Creative Commons photographs from Wikimedia Commons. Production indexing uses frames extracted from the client's selected video.

Explore the Brimtek Workbench

A unified platform for high-quality human data, model evaluation, and multimedia production workflows.

Open Workbench →
BRIMTEK WorkbenchVideo labeling, audio review, subtitling and evaluation in one production workspace
ProjectsVideo LabelingSpeech / AudioSubtitle & DubEvaluate ModelsQA

Video labeling and clipping frame review, scene tags, clip actions

Video labeling preview frame Video Labeling
Clip #08 · market activity · vendor interaction · multilingual review ready
00:12.10 - 00:19.70Scene tagsmarket, produce, vendor, crowd, commerce, street context
00:19.70 - 00:24.20QC tasksverify OCR, approve clip boundary, assign reviewer, export JSON
Reviewer queueAssigned tabsLabel Video, Metadata & Speakers, QA & Agreement

Speech / Audio labeling waveform, segments, translation and dubbing

SPK A
Made in Cambodia is a story of people, craft and opportunity.
00:00.00
Arabic
صنع في كمبوديا هو قصة الناس والحرف والفرص.
aligned
Dub line
Speaker-matched line ready for timing review and TTS synthesis.
00:05.40
TranscribeDiarizeRefineTranslateTTSExport SRT/VTT

Evaluation and QA benchmarks, credits, quality tracking

WER8.4%Arabic broadcast sample
Diarization91.2speaker assignment score
Turnaround14mfrom upload to review pack

Teams can move between labeling, subtitling, indexing, live captioning, and model evaluation while keeping the same project context and review history.

◉ Speech Labeling
▣ Label Images
▧ Label Video
⌁ Evaluate Models
▶ Subtitle & Dub
CC Live CC
◍ Voices & Consent
✓ QA & Agreement
Open BenchmarksSelected multilingual, human-reviewed evaluation samples for testing ASR, diarization, timing, and related models.
Explore benchmarks →
Built in DubaiBRIMTEK operates from DMCC, Dubai, serving multilingual AI, media, call-center, conversational, and security workflows.