Quick Answer
Last verified:
High confidence

Google Cloud Speech-to-Text costs Free to $0.08 per minute as of July 2026, with 14 plans available including a free tier. Plans: Speech-to-Text V2 API - Standard Recognition - 0 to 500,000 minutes (free), Speech-to-Text V2 API - Standard Recognition - 500,000 to 1,000,000 minutes (free), Speech-to-Text V2 API - Standard Recognition - 1,000,000 to 2,000,000 minutes (free), Speech-to-Text V2 API - Standard Recognition - 2,000,000+ minutes (free), Speech-to-Text V2 API - Standard Dynamic Batch Recognition (free), Speech-to-Text V1 API - Standard Recognition with Data Logging - 0 to 60 minutes (free), Speech-to-Text V1 API - Standard Recognition with Data Logging - 60+ minutes (free), Speech-to-Text V1 API - Standard Recognition without Data Logging - 0 to 60 minutes (free), Speech-to-Text V1 API - Standard Recognition without Data Logging - 60+ minutes (free), Medical Dictation - 0 to 60 minutes (free), Medical Dictation - 60+ minutes (free), Medical Conversation - 0 to 60 minutes (free), Medical Conversation - 60+ minutes (free), and Large Workloads / Custom Quote (free). Enterprise pricing is available on request. Pricing depends on your chosen tier, contract length, and negotiated discounts.

Use the interactive pricing calculator to estimate your exact cost based on team size and requirements.

  • Free tier: Yes

Google Cloud Speech-to-Text offers 14 pricing tiers: Speech-to-Text V2 API - Standard Recognition - 0 to 500,000 minutes, Speech-to-Text V2 API - Standard Recognition - 500,000 to 1,000,000 minutes, Speech-to-Text V2 API - Standard Recognition - 1,000,000 to 2,000,000 minutes, Speech-to-Text V2 API - Standard Recognition - 2,000,000+ minutes, Speech-to-Text V2 API - Standard Dynamic Batch Recognition, Speech-to-Text V1 API - Standard Recognition with Data Logging - 0 to 60 minutes, Speech-to-Text V1 API - Standard Recognition with Data Logging - 60+ minutes, Speech-to-Text V1 API - Standard Recognition without Data Logging - 0 to 60 minutes, Speech-to-Text V1 API - Standard Recognition without Data Logging - 60+ minutes, Medical Dictation - 0 to 60 minutes, Medical Dictation - 60+ minutes, Medical Conversation - 0 to 60 minutes, Medical Conversation - 60+ minutes, Large Workloads / Custom Quote. The Speech-to-Text V2 API - Standard Recognition - 500,000 to 1,000,000 minutes plan is higher-volume speech-to-text v2 standard recognition workloads.

Google Cloud Speech-to-Text lists $0-$0.078/minute, but hidden costs like implementation and support add to the total as of July 2026. Key hidden costs: gcp ecosystem overhead doubles or triples effective costs: a production transcription pipeline requires cloud storage ($0.020/gb/month), cloud functions ($0.40/million invocations), pub/sub messaging ($0.40/million messages), and egress fees ($0.08-$0.23/gb) -- budget an additional $50-$300/month in supporting gcp service costs beyond the $0.016/min transcription fee, multi-channel audio is billed per channel: a stereo (2-channel) audio file is billed as 2x the audio duration -- a 60-minute stereo recording costs $1.92 (2 x 60 x $0.016) instead of $0.96, effectively doubling the per-minute rate for multi-channel content like phone calls, enhanced model pricing applies per 15-second increment: audio is billed in 15-second increments rounded up, so a 61-second file costs the same as a 75-second file -- for short audio clips this rounding can increase effective costs by 10-25%. Verified from 5 sources by CostBench.

Hidden Costs Breakdown

1

GCP ecosystem overhead doubles or triples effective costs: A production transcription pipeline requires Cloud Storage ($0.020/GB/month), Cloud Functions ($0.40/million invocations), Pub/Sub messaging ($0.40/million messages), and egress fees ($0.08-$0.23/GB) -- budget an additional $50-$300/month in supporting GCP service costs beyond the $0.016/min transcription fee

2

Multi-channel audio is billed per channel: A stereo (2-channel) audio file is billed as 2x the audio duration -- a 60-minute stereo recording costs $1.92 (2 x 60 x $0.016) instead of $0.96, effectively doubling the per-minute rate for multi-channel content like phone calls

3

Enhanced model pricing applies per 15-second increment: Audio is billed in 15-second increments rounded up, so a 61-second file costs the same as a 75-second file -- for short audio clips this rounding can increase effective costs by 10-25%

4

Data logging opt-in affects pricing: Google may offer lower pricing if you allow your audio data to be used for model improvement (data logging) -- opting out of data logging to maintain privacy may result in higher per-minute rates or reduced access to volume discounts

5

V1 to V2 API migration costs: The V2 API offers Chirp and newer models not available in V1, but migrating requires code changes, testing, and potentially re-architecting audio pipelines -- budget $2,000-$5,000 in engineering time for V1 to V2 migration

6

Egress and cross-region transfer fees apply to transcription results: Downloading transcription results outside of GCP or across regions incurs network egress charges of $0.08-$0.23/GB -- for large-scale operations generating gigabytes of transcript data monthly, this adds $20-$100/month in hidden network costs

Check current Google Cloud Speech-to-Text pricing

Prices and terms change; verify against the live pricing page.

See Google Cloud Speech-to-Text Pricing