The article documents a developer’s exploration of Google Gemini 3.5 Transcribe, focusing on its speech-to-text capabilities, including word-level timestamps, speaker diarization, and a 1,000-item custom vocabulary. It compares transcription behavior to prior models, discusses integration steps (Files API, interaction with TTS, and REST responses), and reveals practical challenges in streaming vs. full transcription, parsing outputs, and UI state handling during mic permission prompts. The author also shares lessons from testing in a development environment, highlighting how transcription results and indexing are affected by UTF-8 byte positions and the distinction between output_text and steps[].content[] in the REST response. The piece emphasizes the importance of validating API responses before going live and notes that auto-correct features can be detrimental for shadowing accuracy, since the goal is to reflect exactly what the user said. Overall, the post offers a practical, engineering-focused look at deploying Gemini 3.5 Transcribe within a language-learning workflow and outlines considerations for future improvements.