Homechevron_rightNewschevron_rightTopicschevron_rightThe deep learning technology "Transformer" has been implemented in the voice recognition API "AmiVoice® API." This has achieved an error improvement rate of up to 17%, significantly improving the recognition rate.

The deep learning technology "Transformer" has been implemented in the voice recognition API "AmiVoice® API." This has achieved an error improvement rate of up to 17%, significantly improving the recognition rate.

On December 12th, we implemented the deep learning technology "Transformer" into almost all of the voice recognition engines of the voice recognition API "AmiVoice API" provided on the voice tech platform for developers "AmiVoice Cloud Platform."

This has resulted in an error reduction rate of up to 17% (according to our research), and a significant improvement in the recognition rate, especially for natural speech.


https://acp.amivoice.com/amivoice_api/

"Transformer" is one of the emerging deep learning technologies.

The recurrent neural network technologies implemented in the conventional voice recognition engine AmiVoice, such as "LSTM (Long Short-Term Memory)" and "Bi-LSTM (Bidirectional Long Short-Term Memory)," incorporate past and future information in the form of memory and calculate present information. However, this memory has the problem that it is difficult to retain information from distant points in time.

In contrast, Transformer performs calculations by directly incorporating information from past and future points in time into the current information, making it possible to effectively use information from distant points in a long input, achieving an even higher recognition rate.

We have now implemented "Transformer" in almost all of the speech recognition engines in "AmiVoice API". Compared to speech recognition engines that implement "Bi-LSTM", the error rate has improved by up to 17% in real-time recognition and up to 13% in batch recognition, significantly improving the recognition rate.

It can be used with the entire lineup of "AmiVoice API" (synchronous HTTP speech recognition API, asynchronous HTTP speech recognition API, WebSocket speech recognition API).


[Speech recognition engine that implements "Transformer"]



General purpose


General-purpose conversation engine, general-purpose voice input engine




For medical use




Medical conversation engine, Medical voice input engine, Pharmaceutical conversation engine, Pharmaceutical voice input engine


For finance and insurance


Finance_conversation engine, Finance_voice input engine, Insurance_conversation engine, Insurance_voice input engine


*The engines for Electronic Medical Record_Voice Input, Chinese (8kHz/16kHz), and English (8kHz) have not been updated to Transformer. They will be updated from time to time in the future.


Features of “AmiVoice API”


1. No.1 voice recognition market share


*


. Convert natural spoken words into text with high accuracy
You can use AmiVoice, a high-precision and high-speed AI voice recognition system with over 25 years of accumulated know-how and data, right from the site. All speech recognition engines and emotion analysis options can be used for free for up to 60 minutes each month.


Five. High quality voice recognition available at low cost

Pay-as-you-go billing based only on the amount of time spoken, not the amount of time recorded. Billing units are not rounded up to the nearest second. You can use a high-quality speech recognition engine at the lowest price in the industry.


3. Speech recognition experts provide free development support

We do everything in-house, from developing voice recognition engines to providing services. Our technical staff will provide direct support free of charge for any technical inquiries you may have, such as during the introduction of the API or individual troubles related to the API after the start of operation.


Four. Achieving high recognition rates with engines that can be selected according to industry and application

In addition to "general-purpose engines" that can be used in a variety of situations, we also have engines specialized for specialized and industry terminology, such as those used in the medical field. Recognition rates can be greatly improved by selecting the engine according to the usage scenario.
By using the dictionary registration function, it is possible to convert in-house terms and proper nouns into text with high accuracy.


5. All service development and operation is done domestically. Available in a secure environment

"AmiVoice API" is developed and operated in Japan. You can use our service with peace of mind as your voice data will not be sent overseas.




(AmiVoice Cloud Platform)

Website

https://acp.amivoice.com/


*Source: ecarlate LLC “Speech Recognition Market Trends 2023” Speech Recognition Software/Cloud Service Market

Inquiries regarding this matter

Management Promotion Headquarters Public Relations Team

Inquiries regarding the contents of this report

Japan's No.1 in Market ShareJapan's No.1 in Market ShareAmiVoiceⓇAmiVoiceⓇ

*Source: ecarlate LLCSpeech Recognition Market Trends 2026'
Speech recognition software/cloud service market

Write with your voice, move with your voice.
AI voice recognition AmiVoice
In various business situations,
This is a technology that enables natural communication between people and machines.