Amazon Polly turns text into lifelike speech, allowing users to create applications that talk, and build new categories of speech-enabled products. Polly's Text-to-Speech (TTS) service uses deep learning technologies to synthesize natural sounding human speech. With dozens of lifelike voices across a broad set of languages, users can build speech-enabled applications that work in different countries.
$4
Per Request
Google Cloud Speech-to-Text
Score 8.3 out of 10
N/A
Speech-to-Text on Google Cloud is a tool used to convert speech into text using an API powered by Google’s AI technologies. The vendor states users can transcribe content in real time or from stored files; deliver a better user experience in products through voice commands; and, gain insights from customer interactions to improve service.
$0.02
per min
Pricing
Amazon Polly
Google Cloud Speech-to-Text
Editions & Modules
Up to 1,000 Requests
$4.00
Per Request
Up to 10,000 Requests
$4.00
Per Request
Speech-to-Text V2 API
$0.016
per min
Speech-to-Text V1 API
$0.024
per min
Offerings
Pricing Offerings
Amazon Polly
Google Cloud Speech-to-Text
Free Trial
No
Yes
Free/Freemium Version
No
Yes
Premium Consulting/Integration Services
No
No
Entry-level Setup Fee
No setup fee
No setup fee
Additional Details
—
Speech-to-Text V1 API
V1 offers data residency for multi region only. Models include short, long, phone call, and video. V1 does not include audit logging. New customers get $300 in free credits and 60 minutes for transcribing and analyzing audio free per month, not charged against your credits.
Speech-to-Text V2 API
V2 offers data residency for multi and single region. Models include short, long, telephony, video, and Chirp. V2 does include audit logging and support for customer managed encryption keys.
More Pricing Information
Community Pulse
Amazon Polly
Google Cloud Speech-to-Text
Considered Both Products
Amazon Polly
No answer on this topic
Google Cloud Speech-to-Text
Verified User
Anonymous
Chose Google Cloud Speech-to-Text
Google Cloud Speech-to-Text outperformed its competitors significantly in terms of accuracy, surpassing any other product available. Additionally, its support for multiple languages was unrivaled in the market. Moreover, for clients with robust bandwidth, Google Cloud …
In our real time meetings or webinars where larger audience are expected we have enabled the captions options with Google Cloud Speech-to-Text tool this start transcribing the complete audio conversation in the neat text format. Also while performing the interview process as well we use this tool to make sure that we adhere to certain rules and are being checked by the superior management team to make sure the transcription has required questions being asked on for quality analysis. Also during the customer call we use this tool to make sure two way communication is transcribed and will be later reviewed when there is an escalation by the superior management
The reasoning behind my 10 is that the UI is very intuitive; I didn't require any formal training to use it. Google's speech-to-text is not just a conversion tool; it helps automate mundane tasks, saves time, and has an almost human-like understanding.
It delivered high accuracy in accented and noisy environments. Regarding its language support, it offers a variety of languages and dialects. Its's Api's are well-documented and easily integrated with our GCP-based stack. Also, its deployment is fast, and it is cost-effective. And if we talk about its translation, it gives real-time and generic translation with great punctuation. Finally, its speaker diarization makes it a cool and yet powerful tool that helps people.