Catalog of Generative AI

Generative AI Catalog

A guide to the generative AI services people are talking about right now, organized by the company that makes them. It covers five areas: text, images, video, voice and music. For each one you can see who makes it, what it is good at, and which version is the newest.

Updated 2026-09-24 Covers 5 areas / 33 services
New to AI? Start here About a 2-minute read
There are a lot of names on this page. Here is how to read it without getting lost.

If you follow the news, new names seem to appear every week, and trying to remember all of them is tiring. You do not need to understand everything at once. Here are the few things that make the rest of the page easier to read.

1 / The company name and the AI name are different things

The names feel tangled because there are three separate layers: the company, the service, and the model.

For example, "the company Anthropic offers a service called Claude, and inside it there is a model called Opus 5." Think of the company as the car maker, the service as the car line, and the model as the trim level.

This page puts one card per company. Normally you only see the name and a short description. Press "Open" to see the list of models inside. Just open the ones that interest you.

2 / Every company has good, better and best

Each company offers everything from smart-but-expensive models to decent, fast and cheap ones. On this page they are labeled Top Mid Light, and the higher a model is in the list, the more capable it is.

For everyday use, a mid-tier model is usually enough. The top tier is for difficult research or complex programming. It also costs several times to more than ten times as much, so there is no need to start with the top model.

3 / The dates tell you what is new

This field changes generation every few months. Each model name has a date next to it, so you can tell whether it is recent or a generation behind.

A model name from more than six months ago may already be discontinued. Old information is easy to run into when you search online, so it helps to check the date.

4 / Common terms

If you run into a word you do not know, you can always come back here.

  • Term
    Model
    The AI itself. The same service can become smarter or cheaper when the model inside is swapped. It is like the engine of a car.
  • Term
    Token
    The unit an AI uses to count text. In English, one token is roughly three quarters of a word. You pay by the amount, so the more text you send in or ask for, the more it costs.
  • Term
    Context (context window)
    How much the AI can read at once. "1 million tokens" is enough to hand over several books together. If you work with long documents, pick a model with a large one.
  • Term
    API
    A way to call the AI directly from a program instead of a chat window. It is used for automation and large volumes of work, and you pay for what you use.
  • Term
    Open weights
    The model itself is published, so you can run it on your own computer or server. Your text never has to leave your machine, which is why it is chosen for work involving information that cannot go outside.
  • Term
    Parameters
    A number describing the size of a model. More usually means smarter, but it also needs bigger hardware to run. You will see it written like "2.8 trillion".
  • Term
    Multimodal
    Able to handle images, audio and video as well as text. Most recent models are moving in this direction.

If you are not sure where to begin: try ChatGPT, Claude or Gemini within their free plans. These three cover a wide range of uses and handle many languages well. Once you know what you want to do next, such as making images or video, the sections below will still be here for you.

Text

12 services

Chatting, writing, translating, and building programs: this is the center of what people mean by "talking to AI." In 2026, US and Chinese companies are competing hard on price, and models with similar ability sometimes appear at one tenth of the price.

If you just want to try something, pick ChatGPT, Claude or Gemini. You can look at the rest once you know what you want to do.

ChatGPT OpenAI / USA
The service that gave ordinary people their first real contact with AI at the end of 2022. It still has the most users, and it is the safe default when in doubt.

The release of ChatGPT in November 2022 is where today's generative AI boom began. The technology already existed, but turning it into "just type into a chat window" is what made it spread so quickly.

The main line now is the GPT-6 series. The top model, Astra, arrived on 3 September 2026, and Sol and Luna followed on 22 September. The previous GPT-5.6 series (released 9 July) has no shutdown date and remains available. The names follow the sun (Sol), earth (Terra) and moon (Luna), so their order matches their size. It is a nicely memorable naming scheme.

Current models (higher = more capable)
  • Top
    GPT-6 Astra2026-09-03
    The flagship at the very top. It is priced high, in the same band as Claude Fable 5.1. You can also choose a "Fast mode" that prioritizes speed, but it costs twice as much again.
  • Upper
    GPT-6 Sol2026-09-22New
    The everyday workhorse. About as smart as the previous GPT-5.6 Sol, at half the price. This is the regular price, not a limited-time offer. On some very hard programming and screen-operation tests, though, the older Sol still has the higher top score.
  • Light
    GPT-6 Luna2026-09-22New
    The fastest and cheapest. Roughly half the price of the previous Luna. In ChatGPT, Free and Go plan users can use it through the desktop app.
  • Previous
    GPT-5.6 Sol / Terra / Luna2026-07-09
    Still available, but if you are starting fresh, GPT-6 does the same job for less. If something already runs on 5.6, there is no need to rush a switch.
  • Images
    GPT Image 22026-04-21
    The successor to DALL-E 3. One of the few models that renders text inside images accurately, including Japanese.

The video service Sora 2 has ended. The app and website closed on 26 April 2026, and the developer API closes on 24 September. It was a much-talked-about service, but OpenAI no longer offers video generation. (This page previously said Sora ended on 24 March. That was the date the shutdown was announced. We apologize for the error.)

Claude Anthropic / USA
Well regarded for writing and programming. In 2026 it became widely known through a developer tool called Claude Code.

The first Claude was released on 14 March 2023, so it actually has more than three years of history. At first only approved users could access it; it opened to everyone with Claude 2 in July of that year. The company was founded by people who had worked at OpenAI, and it has consistently put safety at the front.

The turning point was Claude Code (released February 2025), a tool that lets you ask for development work in plain language from the terminal. It spread rapidly from late 2025 into early 2026, and one survey found three out of four startups using it in 2026. It has also become common to see non-programmers use it to "just try building something that works."

Current models (higher = more capable)
  • Top
    Claude Fable 5.12026-09-01New
    The top model. The base price did not change from the previous version, but the cost of re-reading the same material dropped to a quarter. If you hand it long documents and ask many questions about them, your actual bill will fall considerably.
  • Top
    Claude Mythos 5.12026-09-01Restricted
    A special version of Fable 5.1 with some safety restrictions removed. You cannot simply sign up for it; only vetted organizations have access, for work such as cyber defense or medical research where the usual restrictions would get in the way. If you see the name, you can assume it is not available to the general public.
  • Upper
    Claude Opus 5.52026-09-22New
    The highest model for real-world work. It is 20% cheaper than Opus 5, and Anthropic says it beats Fable 5.1, which costs 2.5 times as much, on programming evaluations. The cost of reading long material has also dropped. Sonnet 5.5 and Haiku 5.5 are announced to follow within weeks.
  • Previous
    Claude Opus 52026-07-24
    The version before Opus 5.5. Its speed was improved on 12 August. When moving to 5.5, some settings (such as turning thinking off) no longer work as before, so take care if you call it from a program.
  • Mid
    Claude Sonnet 52026-06-30
    The everyday workhorse. A price increase was planned for September, but on 10 August it was decided to keep the price unchanged.
  • Light
    Claude Haiku 4.52025-10
    Built for speed.
Gemini Google / USA
Its strength is that it connects directly to Search, Android and Gmail. Veo for video and Nano Banana for images are also Google, making it the company with the widest range.

It started as Bard in March 2023, and the first public demo gave a wrong answer that knocked about $100 billion off the parent company's market value. A rough start. Google then merged Google Brain and DeepMind in April, launched Gemini 1.0 in December, and renamed Bard to Gemini in February 2024.

It was designed from the start to learn from text, images, audio and video together, anticipating today's "you can give it anything" trend. Because it is built into Search, Android and Gmail, many people use it without realizing.

Current models (higher = more capable)
  • Top
    Gemini 3.1 Pro2026
    The flagship for careful reasoning. Its successor, 3.5 Pro, was announced in May but has been delayed three times and was still not out at the end of August.
  • Mid
    Gemini 3.8 Flash2026-09-02New
    Exactly the same price as the previous 3.7 Flash, and better on every published test. An unusual update where only the insides improved. The introductory price lasts until the end of 2026 and doubles on 1 January 2027.
  • Light
    Gemini 3.5 Flash-Lite2026-07-21
    For handling lots of requests quickly.
  • Images
    Nano Banana Pro / 22025-11 / 2026-02-26
    Praised for realistic people and skillful placement of text. Text in Japanese also tends to hold up. Version 2 keeps the quality at one third of the price.
  • Video
    Veo 3.1 / Gemini Omni2026 / 2026-06-30
    Veo was early to produce video with sound. Omni is a new line that handles text, images and video in one model, but its pay-as-you-go price has not been published yet.
Grok xAI (SpaceXAI) / USA
Elon Musk's AI. Its unique feature is reading posts on X (formerly Twitter) as they happen, which makes it strong on current events.

Musk founded the company in March 2023 and released Grok that November. Its team came from OpenAI, DeepMind and Google, and what sets it apart is that it was designed together with X from the start, so it can answer with today's events and trends in mind.

In February 2026 SpaceX acquired xAI, and it is now a division of SpaceX called "SpaceXAI".

Current models (higher = more capable)
  • Top
    Grok 4.72026-09-21New
    The current flagship. Exactly the same price as 4.6, but built on a larger base model and trained further for long tasks that take hours. A "Fast" version at twice the price exists, but it is only available through certain tools.
  • Previous
    Grok 4.62026-08-12
    The flagship before 4.7. Suited to long-running automation, programming and on-screen tasks, with four selectable levels of reasoning depth. External benchmarks score it on par with OpenAI's GPT-5.6 Sol.
  • Upper
    Grok 4.52026-07-09
    Good at remembering long text and using tools. Designed for large programs and for handling big volumes of PDFs and email.
  • Mid
    Grok 4.32026-04-30
    For handling long text cheaply. It is the only one with a discount for bulk processing.
  • Images/Video
    Grok Imagine2026
    Makes both images and video. One of the few that clearly states a bulk discount for video: half price.

In May 2026, Grok 4 and the Grok 3 family were retired all at once. Generations change very quickly, so information about older model names is often unreliable. When reading articles online, check the date.

Muse Meta (Meta Superintelligence Labs) / USA
The company behind Facebook and Instagram. After an earlier stumble it rebuilt its AI team and entered the paid market in earnest in 2026. Low prices are its weapon.

For a long time Meta kept its presence through Llama, a family of free, openly released models. But with Llama 4 in 2025 it came out that an unreleased special version had been used for leaderboard testing, and trust suffered.

In response, Meta created Meta Superintelligence Labs in summer 2025, bringing in Scale AI founder Alexandr Wang as Chief AI Officer and investing $14.3 billion in his company. On 8 April 2026 it released its first model, Muse Spark, shifting from free releases to paid ones.

Current models
  • Top
    Muse Spark 1.32026-09-02New
    Focused on programming. Even cheaper than 1.2, by a wide margin. The strategy of standing out on price is very clear.
  • Upper
    Muse Spark 1.22026-08-05
    According to Mark Zuckerberg, "roughly 25% of the price of equivalent OpenAI or Anthropic models." There is also a cheaper price for customers who allow their data to be used for training.
  • Tool
    Muse Code2026-08-05
    A development assistant that runs in the terminal, clearly aimed at Claude Code.
  • Open
    Muse Glimmer 30B2026-08Open weights
    Released under a permissive license. Meta has also said it will publish the weights of Muse Spark 1.2.
Kimi Moonshot AI / Beijing, China
Highly rated for handling long text and for programming. The latest, K3, is the largest open-weight model ever released.

A Beijing lab that first made its name on how much text it could handle at once, and later earned a reputation for programming too. It is counted among the "four little dragons," China's top four independent AI companies.

Current models
  • Top
    Kimi K32026-07-16Open weights
    At 2.8 trillion parameters, the largest open-weight model ever released. It does not use all of them every time; it activates only the parts each question needs, so it runs lighter than its size suggests. It can read 1 million tokens of text and accepts images and video. The weights were published on 27 July under a custom license, so check the terms before using it for work.
  • Older
    Kimi K2.62026-04-20
    Much cheaper than K3, and the only one that supports the bulk-processing discount.
DeepSeek DeepSeek / Hangzhou, China
The company that embodies "compete on price." In 2025 its low-cost models overturned the industry's assumptions.

In 2025 it released models with similar ability at a fraction of the price, making everyone ask whether AI really had to cost so much. It was said to have upended the calculations of investors and cloud companies. Its focus on efficiency has been consistent.

However, on 16 August 2026 it moved to time-of-day pricing, which amounted to a price increase. Part of a May price cut that had been described as "permanent" was reversed after three months. Even a company known for being cheap can do this.

Current models
  • Top
    DeepSeek V4 Pro2026-04-24
    A top-tier model priced like other companies' light models.
  • Special
    DeepSeek R12025
    Specialized for careful reasoning. It kicked off this whole category.
  • Light
    DeepSeek V4.1-Flash2026-09-10
    The newest. It can now read images, and its weights are openly available. At one point DeepSeek announced it would also take over requests for the top model from 14 September, but that plan was cancelled. The top model remains available as before.
  • Light
    DeepSeek V4 Flash2026
    The previous budget model. It has been retired: if you request it by this name, the new V4.1-Flash actually answers (and you pay the new model's price). The image-reading Flash Vision (experimental), added on 21 August, is handled the same way.

Prices change by time of day. Peak hours are weekdays 01:00–04:00 and 06:00–10:00 UTC (10:00–13:00 and 15:00–19:00 in Japan). Running batch jobs outside those hours costs half as much.

GLM Zhipu AI (Z.ai) / Beijing, China
A company born out of Tsinghua University that listed in Hong Kong in January 2026. Its permissive licenses make it easy for businesses to use.

Spun out of Tsinghua University and known internationally as Z.ai. It listed on the Hong Kong Stock Exchange in January 2026 with a market value of about $6.6 billion, giving it an unusually solid financial base for a Chinese AI company.

Its hallmarks are releases under permissive licenses and a focus on automation and programming. The combination of research depth and easy licensing makes it a frequent pick for enterprise adoption.

Current models
  • Top
    GLM-5.22026-06-13Open weights
    744 billion parameters, with a 1-million-token context. The weights were published within days of the announcement.
  • Mid
    GLM-5-Turbo2026-03-15
    A separate line built for speed.
Qwen Alibaba / China
Alibaba's AI. The second most used chat service in China, run hand in hand with Alibaba's cloud business.

An e-commerce and cloud giant that treats AI as the front door to its own cloud. It has 166 million monthly users, second in China, and even designs its own chips, aiming to own the whole stack.

The video model Wan is also from Alibaba. It drew attention in August 2026 with Wan 3.0, which can make video from documents.

Current models
  • Top
    Qwen3.8-Max2026-08-03
    The flagship you can actually use today. Qwen has long been a strong presence among openly released models too.
  • Announced
    Qwen 4 Max and 3 othersannounced 2026-09-22
    At a conference on 22 September, four models were previewed: Max, Plus, Flash and 27B (for local use). No prices or release dates have been announced yet. For now, it is only a preview.
  • Video
    Wan 3.02026-08-24New
    See "Video" below.
MiniMax MiniMax / Shanghai, China
Text, voice, video, music and images all under one account. Best known for its Hailuo video models.

Founded in 2021 and listed in Hong Kong; another of the "four little dragons." What sets it apart is that it runs both the M line for text and the Hailuo line for video and voice, all usable from one account. If you want to try many things, having them in one place is convenient.

Current models
  • Text
    MiniMax-M32026
    Accepts up to 512K tokens of input. Offered at a permanently 50%-off price.
  • Video
    Hailuo H32026-07-31
    See "Video" below.
  • Voice
    MiniMax Speech2026
    See "Voice" below.
Doubao / Seed ByteDance / China
The company behind TikTok. Doubao is China's most used chat service with 345 million monthly users, and ByteDance is also at the front in video and images.

Its consumer service Doubao has 345 million monthly users, the largest in China. The research division, ByteDance Seed, makes Seedance for video and Seedream for images.

In video it is considered the current quality leader. Seedance 2.0 overtook Google Veo 3, OpenAI Sora 2 and Runway Gen-4.5 in video rankings.

Current models
  • Video
    Seedance 2.5 / 2.02026-07-31 / 2026-02
    See "Video" below.
  • Images
    Seedream 5.0 Pro2026-08-12
    See "Images" below.
Fugu Sakana AI / Tokyo, Japan
A Japanese company with an unusual approach: instead of answering everything itself, it routes each job to whichever other AI is best at it.

Founded in Tokyo in 2023, it is known for having among its founders a researcher who co-wrote at Google the paper that became the foundation of today's AI (the Transformer). The name Sakana ("fish" in Japanese) comes from the idea of many small things working together like a school of fish.

Fugu, released on 11 September 2026, is not a single big AI. It decides which AI is good at a given job, sends the work to other companies' models, and then combines the answers. It is a coordinator (orchestrator). Rather than building a huge model from scratch, the idea is to get stronger and cheaper by using existing models well.

Current models
  • Top
    Fugu Ultra v2.02026-09-11
    Aims for the best possible answer on difficult, many-step work.
  • Mid
    Fugu Max2026-09-11
    Aims for the best result per dollar. Sakana says it is 40–60% cheaper than models doing similar work.

It is a significant step for a Japanese company to appear in this list with this level of performance and price. Keep in mind, though, that it works by coordinating other companies' AIs, which is a little different from "an AI made in Japan" in the usual sense.

Images

6 services

Making pictures, illustrations and photo-like images. One thing that matters a lot in practice is whether text inside the image comes out correctly, and companies differ widely here. For posters and slides, you will probably choose by how accurately text is reproduced rather than by how pretty the picture is.

For images with text, GPT Image 2.5 or Nano Banana Pro. For purely beautiful images, Midjourney. These two goals are quite different.

GPT Image 2.5 / 2 OpenAI / USA
A leader in rendering text correctly, including Japanese. Works well for summary images and diagrams.

Released on 21 April 2026 as the successor to DALL-E 3, it also powers ChatGPT's image feature. It reproduces text accurately, even in Japanese, so it is a first choice when you want words inside your images.

You can choose low, medium or high quality, and the price differs by about 30 times. Draft in low quality and use high quality only for the final version to save a lot.

GPT Image 2.5 was released on 8 September 2026. It comes in two versions: Flare, built for speed and volume, and Sunburst, built for precise editing of your own images. Both cost the same. It is considerably more expensive than version 2, though, so version 2 is fine until you know what you need.

Nano Banana Pro / 2 Google / USA
Known for realistic people and skillful text placement. Version 2 keeps the quality at one third of the price.

Officially part of the Gemini Image family, but far better known by its nickname, "Nano Banana." Pro came out in November 2025, and the faster, cheaper Nano Banana 2 followed on 26 February 2026.

Pro is good at realistic people and at layering text, and holds up relatively well even on Japanese posters and social media images. Version 2 supports 4K while producing images in 3–5 seconds, which suits high-volume work.

Midjourney Midjourney / USA
Long loved for artistic beauty. Better for making artwork than for making slides.

From the beginning it has won fans on one point: the images are simply beautiful. It is chosen by people who value artistic quality over technical accuracy.

V8 Alpha arrived on 17 March 2026, followed by V8.1 in April, which became the default in June. Rendering is five times faster, and 2K output is supported.

Seedream 5.0 ByteDance Seed / China
It can edit as well as create, and it can blend several reference images, which is useful in real production work.

Released on 12 August 2026. Its features are that it can fix images as well as make them and that it can combine multiple reference images, which fits professional production workflows. The first reference image is free.

FLUX.2 Black Forest Labs / Germany
Founded by former Stable Diffusion developers. A good balance of quality and ease of use.

The FLUX.2 family arrived on 25 November 2025, taking over as flagship from FLUX.1. On 15 January 2026 a freely usable lightweight version, Klein 4B, followed. It is popular with people who want to run models on their own machines.

Stable Diffusion Stability AI / UK
The model that opened image generation to everyone. Its biggest strength is that you can run it on your own computer.

When it was released in 2022, it let anyone run image generation on their own computer, and that spread the field enormously. The culture of derivative models and fine-tuning also grew from here.

Stable Diffusion 4 arrived on 6 April 2026. It is the first generation since 2022 with a fundamentally new design, moving to the same approach as other companies' latest models.

Video

7 services

The area with the fastest turnover this past year. It is normal for a model name from six months ago to be gone already. In 2026 the competition runs in three directions: making sound together with video, making long clips in one go, and making video from documents.

For quality, Seedance. To try things cheaply, Veo 3.1 Lite. To make video from documents, Wan 3.0. Prices in this area vary by more than 10 times, so start with something cheap to get a feel for it.

Seedance ByteDance Seed / China
Current quality leader. It overtook Veo, Sora and Runway in rankings.

Version 2.0 arrived in February 2026 and took first place in video rankings, ahead of Google Veo 3, OpenAI Sora 2 and Runway Gen-4.5. Version 2.5 followed on 31 July.

With 2.5 you can make up to 30 seconds, including sound, in one pass. You can supply up to 30 reference images, 10 videos and 10 audio clips, and edit down to the second. The official output resolution is up to 720p, however. Resellers advertising 1080p or 4K may simply be upscaling it themselves, so be careful.

Wan 3.0 Alibaba / China
It makes video from PDFs, slides and spreadsheets. No other model accepts input like this.

Released on 24 August 2026. Its main feature is that besides text and images, you can hand it documents, spreadsheets, slides, PDFs and web pages as they are, and it builds a video from them. No other company offers this idea of passing in your existing materials directly.

Clips run 2–30 seconds at 480p, 720p or 1080p. Failed generations are not charged, which is a real help in practice. Since the public beta on 6 August, it has been used for short dramas, ads, tourism promotion and music videos.

Veo / Gemini Omni Google / USA
An early leader in video with sound, with cheap options too. A wide range.

Veo was the first major model to generate dialogue, sound effects and ambient sound together with the picture. The current 3.1 comes in three quality levels, and the top and bottom differ in price by more than 10 times. If you do not need sound, Lite is among the cheapest video options of all.

On 30 June 2026 a new line, Gemini Omni Flash, also appeared. It handles text, images and video in one model, but its pay-as-you-go price has still not been published, which makes it hard to budget for work.

Hailuo H3 MiniMax / China
A unified model that understands text, images, video and audio together, and generates sound at the same time.

Released on 31 July 2026. It is designed to understand four kinds of input in a single model, and accepts up to 9 reference images, 3 videos and 3 audio clips. It produces 5–15 seconds at 2K and 24fps, with stereo sound.

In practice, resolution is 2K only. The price list does show a cheaper 768p tier, but it is invitation-only and cannot actually be selected through the API, so you are billed at the 2K price. Do not judge it as cheap from the price list alone.

Kling Kuaishou / China
Supports 4K at 60fps and lip-sync in multiple languages. Well regarded for film-like work.

Version 3.0 arrived on 4 February 2026. It can output 4K, 60fps, 15 seconds in one go and matches mouth movement in multiple languages. A faster Turbo version was added in mid-June. Pricing uses credits, and consumption depends on resolution and whether sound is included.

Gen-4.5 Runway / USA
A veteran for video professionals, valued for fine control.

This company was in real video production back when AI video could only make a few seconds. Chinese companies are overtaking it on raw quality, but it keeps its fans thanks to fine-grained control for creators.

Ray 3 Luma AI / USA
The price changes a lot with resolution. At 720p you can try it cheaply.

Known as Dream Machine. The price per second is nearly four times higher at 1080p than at 720p, so drafting at 720p and finishing at 1080p works well.

OpenAI's Sora 2 has ended. The app and website closed on 26 April 2026, and the developer API closes on 24 September. (This page previously said Sora ended on 24 March. That was the date it was announced. We have corrected it.) Once almost synonymous with AI video, it is no longer available. Its exit marked the shift of leadership in video toward ByteDance, Google and Alibaba.

Voice (speech and voices)

4 services

Reading text aloud and creating voices themselves. Used for narration, audiobooks, dubbing and game dialogue. Since 2026, more of these tools do not just read aloud but act as someone you can talk to.

To have text read aloud, ElevenLabs. To have a conversation partner, GPT Live. These two serve completely different purposes, so start from what you want to do.

GPT Live OpenAI / USA
The engine behind ChatGPT's voice conversations. It keeps listening while you speak, so you can interrupt and the conversation carries on.

It became available in the ChatGPT app on 8 July 2026, and from 10 September 2026 you can also build it into your own service.

The key difference is that it can listen while it talks. Earlier voice AIs waited for you to finish before replying, and stopped if you talked over them. GPT Live listens even mid-sentence, so you can say "no, that's not what I meant" and the conversation keeps going. The response delay also dropped from 1.41 to 0.798 seconds.

In ChatGPT, Free plan users get a lighter version (mini), and Go, Plus and Pro users get the full version.

If you build it into your own service, read the pricing carefully. The listed 5 cents per minute covers only the listening-and-speaking part. The actual thinking is done by a separate AI, and that is billed separately. As a rough guide, a 10-minute conversation costs about $0.51 with a light model and about $0.86 with a smart one. On top of that, silent time is billed at the same rate.

One more thing: the 12 published voices are English accents plus Brazilian Portuguese and Filipino. No Japanese voice is listed. If you plan to use it in another language, test it before committing.

ElevenLabs ElevenLabs / USA, founded by Poles
The leader in this area, ahead on naturalness and voice-cloning accuracy.

Started in 2022 by two childhood friends from Poland, it reached an $11 billion valuation in four years and raised $500 million in February 2026.

Its text-to-speech and voice cloning are especially well regarded and are widely used for audiobooks, dubbing, games and video production. There are speed-focused and quality-focused options, differing in price by about two times.

MiniMax Speech MiniMax / China
Usable under the same account as MiniMax text, video and music. Good when you want everything in one place.

Two tiers: Turbo for speed and HD for quality. On its own it is priced about the same as ElevenLabs, but because MiniMax covers text, images, video and voice under one account, it becomes attractive if you want to try many things.

Music is a different story now. Since 20 August 2026, the music and lyrics generation APIs are not accepting new sign-ups. Existing users can continue, but new users must either use the MiniMax Audio app or run the published model on their own hardware.

OpenAI TTS / Google TTS OpenAI and Google / USA
Text-to-speech built into the big providers' platforms. Convenient if you already have an account.

Both are offered as part of the main platforms. Compared with the specialist ElevenLabs, they can fall a step behind on expressiveness, but if you already use them for text, you can start without a new contract.

On 26 August 2026 Google also released Gemini 3.5 Transcribe, specialized for transcription, as it builds out its voice lineup.

Then on 15 September came "Gemini 3.8 Live," a voice model you can hold a conversation with. Like GPT Live, you speak to it and it answers by voice. It is reported to support 97 languages, including Japanese, and it is priced low, at roughly 2.3 cents per minute for listening and speaking combined. The language information comes from third-party coverage and has not been confirmed in Google's own documentation, so try it before deciding.

Google's text-to-speech has four price levels depending on voice quality, with about a 40-times difference between top and bottom. The lowest level includes 4 million free characters per month, so small volumes cost nothing.

Music

4 services

Creating whole songs and backing tracks. What makes this area different is that lawsuits over the rights to training data are still ongoing, so check the terms before using it for work.

For quality, Suno. For peace of mind in commercial use, ElevenMusic. ElevenMusic states that its training data is properly licensed.

Suno Suno / USA
Rated highest for output quality, though debate over rights continues.

As of February 2026 it had about 2 million paying users. Its output is highly rated across genres: pop, rock, electronic and ambient.

In this area, even the way services are offered is changing. MiniMax stopped taking new music API sign-ups on 20 August 2026 and moved to an app and an open model. If you plan to use one for work, check not only the price but also whether you will still be able to sign up in the future.

Meanwhile, lawsuits over training data are ongoing, and the conditions for commercial use are not settled. This is something to check and decide for yourself. There is no official API; access is through third-party providers.

ElevenMusic ElevenLabs / USA
The key difference is that its training data is properly licensed, which makes it easier to use for work.

ElevenLabs, which built its name in voice, moved into music in August 2025 and launched a standalone app on 1 April 2026.

The difference is less about quality and more about clear rights. It states that it trained on data 100% licensed from music rights organizations. It avoids the rights issues Suno and Udio face, so it is a strong option for commercial use.

Udio Udio / USA
On par with Suno in quality, and likewise caught up in the rights debate.

Alongside Suno, one of the most highly rated services for output. It faces similar training-data lawsuits, so terms for commercial use are in flux. There is no pay-as-you-go API; only monthly subscriptions.

Lyria Google / USA
Unusual in having a version that keeps playing music in real time.

Google's music model. Besides the standard version that makes one song at a time, there is a real-time version that keeps the music flowing without a break, useful for things like continuous background music during a live stream.