Speech-to-Text on Google Cloud is a tool used to convert speech into text using an API powered by Google’s AI technologies. The vendor states users can transcribe content in real time or from stored files; deliver a better user experience in products through voice commands; and, gain insights from customer interactions to improve service.
$0.02
per min
Phrase Localization Platform
Score 9.2 out of 10
Small Businesses (1-50 employees)
The Phrase Localization Platform is an AI-powered language platform that integrates translation, scoring, and automation tools in one place for businesses and language service providers. It offers scalability, a vendor-neutral approach, and advanced analytics for performance optimization, with single sign-on to facilitate easy setup.
$27
per month
Pricing
Google Cloud Speech-to-Text
Phrase Localization Platform
Editions & Modules
Speech-to-Text V2 API
$0.016
per min
Speech-to-Text V1 API
$0.024
per min
Freelancer
$27
per month
Direct plan: Starter
$135
per month
LSP plan: Professional
$875
per month
Direct plan: Team
1,045
per month
Direct plan: Business
4,395
per month
Direct plan: Enterprise
Custom
per month
LSP plan: Business
Custom
per month
LSP plan: Enterprise
Custom
per month
Offerings
Pricing Offerings
Google Cloud Speech-to-Text
Phrase Localization Platform
Free Trial
Yes
Yes
Free/Freemium Version
Yes
No
Premium Consulting/Integration Services
No
No
Entry-level Setup Fee
No setup fee
Optional
Additional Details
Speech-to-Text V1 API
V1 offers data residency for multi region only. Models include short, long, phone call, and video. V1 does not include audit logging. New customers get $300 in free credits and 60 minutes for transcribing and analyzing audio free per month, not charged against your credits.
Speech-to-Text V2 API
V2 offers data residency for multi and single region. Models include short, long, telephony, video, and Chirp. V2 does include audit logging and support for customer managed encryption keys.
—
More Pricing Information
Community Pulse
Google Cloud Speech-to-Text
Phrase Localization Platform
Considered Both Products
Google Cloud Speech-to-Text
Verified User
Anonymous
Chose Google Cloud Speech-to-Text
Much better and more accurate than integrated Microsoft dictate or translate.
I like Google Cloud Speech-to-Text the most when it comes to other apps I have used so far. It have reduced my work, saved lot of time and made me less stress in meetings. It has also helped us in taking requirement gathering, knowledge transfer important notes to further …
While both Speechify and Google Speech-to-text do the job, certain elements that I find missing on Speechify are: it only works on Desktop with Windows OS, the customizations aspect is missing, there is no mobile app support (people these days want everything on their mobile …
I've also trialed IBM Watson Speech to Text for similar use cases. While both are highly capable, I find the Google Cloud Speech-to-Text software's accuracy and integrations to be a cut above. Harnessing Google's speech recognition prowess has elevated our firm's value …
Google Cloud Speech-to-Text outperformed its competitors significantly in terms of accuracy, surpassing any other product available. Additionally, its support for multiple languages was unrivaled in the market. Moreover, for clients with robust bandwidth, Google Cloud …
1. It's an efficient tool for improving efficiency by saving a lot of time in typing. 2. It saves at least 40-50% of our time, thus increasing efficiency. The amazing thing I liked about it is the accuracy with multiple accents & multiple languages. 3. It also takes …
The accuracy of Google Cloud Speech-to-Text is much better than any other tool. It has better API integration with 3rd party tools. The transcription is on at real-time basis with the best efficiency. It has good language support from across the globe. It provides better noise …
Google Cloud Speech-to-Text is better than these other services. The main driver is the cost for the service and what you get, the value proposition is very good. Also, the scalability of Google Cloud Speech-to-Text is great, so that down the line, as our needs change and …
Google Cloud Speech-to-Text shows an impressive ROI with increased efficiency, time savings, accuracy, speed, productivity, customer satisfaction, and cost-effectiveness.
Dragon is a long stading product in the market but the intuitiveness and that it was part of the Google ecosystem made me switch easily. it also allowed me to easily integrate with other Google product which we are so accustomed to, where Dragon was lacking and did not provide …
Google is far more ahead when compared to Amazon product with similar capabilities and helps to understand and interpret the speech in a much better and clarified way which helps to solve the business use case in a quicker manner and helps to reduce the over all time taken …
Is very easy to implement. We simply selected each feature and obtain great data from meetings. Our team was pleased to use it. It converted data accurately. We recommend it!!
Google Cloud Speech to Text has a significantly cleaner and easier-to-use User Interface. If the user is already familiar with the Google Cloud product suite, then onboarding with this software will be an extremely smooth process. If a user has previously used other …
We use Google Speech to Text on the recommendation of a partner who uses it, in fact we do not evaluate other applications such as Amazon Transcript or similar
They just remind me of each other. Whenever I have a question, whether for personal or for professional reasons, I take out my smartphone, click the Gemini app, and then click the mic to ask my question and have the answer read back to me. I love Googles AI System.
It delivered high accuracy in accented and noisy environments. Regarding its language support, it offers a variety of languages and dialects. Its's Api's are well-documented and easily integrated with our GCP-based stack. Also, its deployment is fast, and it is cost-effective. …
It is very expensive and more difficult to navigate. Compared to the other support is slow and ineffective. Dealing with anything more involved is all but impossible. It has become more prone to bugs since being bought by TMS. Support suggests things like "use incognito" mode …
I did not select Memsource. I use it because my clients use it, but I find it very useful especially due to a friendly interface and quick jumping between segments of different statuses during translations.
SDL Trados Studio is more robust but also much more bloated and resource-intensive and ultimately less flexible. I use both daily, but Memsource is my go-to in almost every situation.
Memsource is quicker and easier to use. You can even start translating on your mobile phone or tablet, as long as you have an internet connection. Plus Memsource shows me a live preview of my translation. I think Memsource is great for beginners as well, since you don't need a …
Memsource is the best, because it does not have any unneeded functionality that would compromise its intuitive use, and those functionalities it includes are all easily accessible.
Memsource stands out because it can be used in ANY browser. Thanks to its mobile app, it is easy to accept and review the work that translators are assigned.
I also use Wordbee, but the thing I don't like it most about it is that the number of segments you can display at a time is limited to 100 and you must move around many pages if the job size is large. I use SDL Trados Studio as well, but not cloud-based, so not really …
Memsource is the easiest to use among all other CAT tools out there. Even when you use it for the first time, it will be easy for anyone to use. You don't even need much instruction or a tutorial process. Still, it allows you to do a lot. It's easy to find some features you …
There is no true comparison because Memsource consistently outperforms the competition because it is simpler to use and has greater TM and MT integration.
Memsource is way cheaper and easier to use. Other softwares seem really old-fashioned and out of date. They are very difficult to use for a new learner and the UI is not good looking. Also, the price for products like Trados is too much expensive for a new freelancer and many …
Memsoucre has a more friendly user interface. It keeps offering new features (like machine translation add on) which wasn't there when we first started using Memsource. Very smooth and easy communication with customer service team. My tickets are addressed quickly and …
I have been using memoQ and SDL Trados Studio for a longer time, but the Web editor for Memsource is the easiest to use of all. I do like the functionality in memoQ better, but Memsource is easier for our translators.
What distinguishes Memsource from other translation tools is that it does not require massive specifications to run. Average computers can run it smoothly without needing 8 or 12 Gb ram. Being able to use the browser to do the translation task is also another excellent feature. …
Memsource is the platform chosen by one of my clients, Booking.com. I have direct contact with my client, and the machine translation service offered is better than the one offered by its competitors. It's more intuitive and easy to use as well.
Memsource includes all the basic features a freelance translator needs in a very user-friendly approach. All you need is a click away and provides quick results.
Memsource has a very good user interface that people can quickly learn and start using. Also, its analysis features are top class which helps me provide a detailed estimate to my client based on the file particulars. It has a QA feature that helps me do a very high-quality …
Memsource is more user friendly and innovative. The interface looks better and more accessible. I have not experienced any downtime with Memsource. The times that I used Memsource it didn't give me any errors or issues. But with the other tool, there were several times when …
In our real time meetings or webinars where larger audience are expected we have enabled the captions options with Google Cloud Speech-to-Text tool this start transcribing the complete audio conversation in the neat text format. Also while performing the interview process as well we use this tool to make sure that we adhere to certain rules and are being checked by the superior management team to make sure the transcription has required questions being asked on for quality analysis. Also during the customer call we use this tool to make sure two way communication is transcribed and will be later reviewed when there is an escalation by the superior management
I like Memsource because when I am not home and I don't have my laptop, I can borrow a computer, log in to my Memsource account and I'm ready to begin translating. I can even download the source files and check the TM and glossary there. It's not necessary to download the software and lose time on that.
In the case of EN to JA translation, we enter text and then convert it to get the correct final text (word or phrase) which uses correct Chinese character(s), as there are often multiple Chinese characters with the same reading but different meanings. When those conversion options are displayed, it is usually possible to change our selection among them by hitting the Tab key, but in Memsource, hitting the Tab key makes us leave the text conversion and move to the source segment, so we must make sure to use arrow keys to choose the right conversion option when using Memsource.
When there are tag elements used in the source text, the tags must exist in the target text of course, but the tag order also must be the same. The tags cannot be moved around in the segment, which causes problems in the case of the Japanese language, because the word order differs between EN and JA.
In the Japanese language, Italics are basically not used, so there must be no texts between the tags which specify Italic font. But then the segment cannot be confirmed and the job status cannot be changed to "Complete".
I use and manage an Academic edition (approx. 15 students) and the corporate license (10 users). In both cases, the features are easy to find, set and use. My students and my translators understand the dynamics of the platform easily and get used to them quickly.
The reasoning behind my 10 is that the UI is very intuitive; I didn't require any formal training to use it. Google's speech-to-text is not just a conversion tool; it helps automate mundane tasks, saves time, and has an almost human-like understanding.
It is designed to be a lean system, but that means that things are clustered in weird groups and a lot of trial and error is involved getting the settings right. If something goes wrong first level support is very quick and has very fast answers which are quite typical of first level support and essentially equate to asking you if you have switched it off and on yet. Anything more involved is challenging for support to address.
We rarely use support, but most questions were answered in a timely fashion, although we didn't exactly find them satisfactory. That's mostly the fault of the software and not the Support team because we asked for things that Memsource couldn't do.
It delivered high accuracy in accented and noisy environments. Regarding its language support, it offers a variety of languages and dialects. Its's Api's are well-documented and easily integrated with our GCP-based stack. Also, its deployment is fast, and it is cost-effective. And if we talk about its translation, it gives real-time and generic translation with great punctuation. Finally, its speaker diarization makes it a cool and yet powerful tool that helps people.
It is very expensive and more difficult to navigate. Compared to the other support is slow and ineffective. Dealing with anything more involved is all but impossible. It has become more prone to bugs since being bought by TMS. Support suggests things like "use incognito" mode to log in. Use a different browser to log-in. That is a concern given the customer data involved if things at that level are not being fixed. If that's what it's like in the showroom, what is the kitchen like?