A guide to the generative AI services people are talking about right now, organized by the company that makes them. It covers five areas: text, images, video, voice and music. For each one you can see who makes it, what it is good at, and which version is the newest.
If you follow the news, new names seem to appear every week, and trying to remember all of them is tiring. You do not need to understand everything at once. Here are the few things that make the rest of the page easier to read.
The names feel tangled because there are three separate layers: the company, the service, and the model.
For example, "the company Anthropic offers a service called Claude, and inside it there is a model called Opus 5." Think of the company as the car maker, the service as the car line, and the model as the trim level.
This page puts one card per company. Normally you only see the name and a short description. Press "Open" to see the list of models inside. Just open the ones that interest you.
Each company offers everything from smart-but-expensive models to decent, fast and cheap ones. On this page they are labeled Top Mid Light, and the higher a model is in the list, the more capable it is.
For everyday use, a mid-tier model is usually enough. The top tier is for difficult research or complex programming. It also costs several times to more than ten times as much, so there is no need to start with the top model.
This field changes generation every few months. Each model name has a date next to it, so you can tell whether it is recent or a generation behind.
A model name from more than six months ago may already be discontinued. Old information is easy to run into when you search online, so it helps to check the date.
If you run into a word you do not know, you can always come back here.
If you are not sure where to begin: try ChatGPT, Claude or Gemini within their free plans. These three cover a wide range of uses and handle many languages well. Once you know what you want to do next, such as making images or video, the sections below will still be here for you.
Chatting, writing, translating, and building programs: this is the center of what people mean by "talking to AI." In 2026, US and Chinese companies are competing hard on price, and models with similar ability sometimes appear at one tenth of the price.
If you just want to try something, pick ChatGPT, Claude or Gemini. You can look at the rest once you know what you want to do.
The release of ChatGPT in November 2022 is where today's generative AI boom began. The technology already existed, but turning it into "just type into a chat window" is what made it spread so quickly.
The main line now is the GPT-6 series. The top model, Astra, arrived on 3 September 2026, and Sol and Luna followed on 22 September. The previous GPT-5.6 series (released 9 July) has no shutdown date and remains available. The names follow the sun (Sol), earth (Terra) and moon (Luna), so their order matches their size. It is a nicely memorable naming scheme.
The video service Sora 2 has ended. The app and website closed on 26 April 2026, and the developer API closes on 24 September. It was a much-talked-about service, but OpenAI no longer offers video generation. (This page previously said Sora ended on 24 March. That was the date the shutdown was announced. We apologize for the error.)
The first Claude was released on 14 March 2023, so it actually has more than three years of history. At first only approved users could access it; it opened to everyone with Claude 2 in July of that year. The company was founded by people who had worked at OpenAI, and it has consistently put safety at the front.
The turning point was Claude Code (released February 2025), a tool that lets you ask for development work in plain language from the terminal. It spread rapidly from late 2025 into early 2026, and one survey found three out of four startups using it in 2026. It has also become common to see non-programmers use it to "just try building something that works."
It started as Bard in March 2023, and the first public demo gave a wrong answer that knocked about $100 billion off the parent company's market value. A rough start. Google then merged Google Brain and DeepMind in April, launched Gemini 1.0 in December, and renamed Bard to Gemini in February 2024.
It was designed from the start to learn from text, images, audio and video together, anticipating today's "you can give it anything" trend. Because it is built into Search, Android and Gmail, many people use it without realizing.
Musk founded the company in March 2023 and released Grok that November. Its team came from OpenAI, DeepMind and Google, and what sets it apart is that it was designed together with X from the start, so it can answer with today's events and trends in mind.
In February 2026 SpaceX acquired xAI, and it is now a division of SpaceX called "SpaceXAI".
In May 2026, Grok 4 and the Grok 3 family were retired all at once. Generations change very quickly, so information about older model names is often unreliable. When reading articles online, check the date.
For a long time Meta kept its presence through Llama, a family of free, openly released models. But with Llama 4 in 2025 it came out that an unreleased special version had been used for leaderboard testing, and trust suffered.
In response, Meta created Meta Superintelligence Labs in summer 2025, bringing in Scale AI founder Alexandr Wang as Chief AI Officer and investing $14.3 billion in his company. On 8 April 2026 it released its first model, Muse Spark, shifting from free releases to paid ones.
A Beijing lab that first made its name on how much text it could handle at once, and later earned a reputation for programming too. It is counted among the "four little dragons," China's top four independent AI companies.
In 2025 it released models with similar ability at a fraction of the price, making everyone ask whether AI really had to cost so much. It was said to have upended the calculations of investors and cloud companies. Its focus on efficiency has been consistent.
However, on 16 August 2026 it moved to time-of-day pricing, which amounted to a price increase. Part of a May price cut that had been described as "permanent" was reversed after three months. Even a company known for being cheap can do this.
Prices change by time of day. Peak hours are weekdays 01:00–04:00 and 06:00–10:00 UTC (10:00–13:00 and 15:00–19:00 in Japan). Running batch jobs outside those hours costs half as much.
Spun out of Tsinghua University and known internationally as Z.ai. It listed on the Hong Kong Stock Exchange in January 2026 with a market value of about $6.6 billion, giving it an unusually solid financial base for a Chinese AI company.
Its hallmarks are releases under permissive licenses and a focus on automation and programming. The combination of research depth and easy licensing makes it a frequent pick for enterprise adoption.
An e-commerce and cloud giant that treats AI as the front door to its own cloud. It has 166 million monthly users, second in China, and even designs its own chips, aiming to own the whole stack.
The video model Wan is also from Alibaba. It drew attention in August 2026 with Wan 3.0, which can make video from documents.
Founded in 2021 and listed in Hong Kong; another of the "four little dragons." What sets it apart is that it runs both the M line for text and the Hailuo line for video and voice, all usable from one account. If you want to try many things, having them in one place is convenient.
Its consumer service Doubao has 345 million monthly users, the largest in China. The research division, ByteDance Seed, makes Seedance for video and Seedream for images.
In video it is considered the current quality leader. Seedance 2.0 overtook Google Veo 3, OpenAI Sora 2 and Runway Gen-4.5 in video rankings.
Founded in Tokyo in 2023, it is known for having among its founders a researcher who co-wrote at Google the paper that became the foundation of today's AI (the Transformer). The name Sakana ("fish" in Japanese) comes from the idea of many small things working together like a school of fish.
Fugu, released on 11 September 2026, is not a single big AI. It decides which AI is good at a given job, sends the work to other companies' models, and then combines the answers. It is a coordinator (orchestrator). Rather than building a huge model from scratch, the idea is to get stronger and cheaper by using existing models well.
It is a significant step for a Japanese company to appear in this list with this level of performance and price. Keep in mind, though, that it works by coordinating other companies' AIs, which is a little different from "an AI made in Japan" in the usual sense.
Making pictures, illustrations and photo-like images. One thing that matters a lot in practice is whether text inside the image comes out correctly, and companies differ widely here. For posters and slides, you will probably choose by how accurately text is reproduced rather than by how pretty the picture is.
For images with text, GPT Image 2.5 or Nano Banana Pro. For purely beautiful images, Midjourney. These two goals are quite different.
Released on 21 April 2026 as the successor to DALL-E 3, it also powers ChatGPT's image feature. It reproduces text accurately, even in Japanese, so it is a first choice when you want words inside your images.
You can choose low, medium or high quality, and the price differs by about 30 times. Draft in low quality and use high quality only for the final version to save a lot.
GPT Image 2.5 was released on 8 September 2026. It comes in two versions: Flare, built for speed and volume, and Sunburst, built for precise editing of your own images. Both cost the same. It is considerably more expensive than version 2, though, so version 2 is fine until you know what you need.
Officially part of the Gemini Image family, but far better known by its nickname, "Nano Banana." Pro came out in November 2025, and the faster, cheaper Nano Banana 2 followed on 26 February 2026.
Pro is good at realistic people and at layering text, and holds up relatively well even on Japanese posters and social media images. Version 2 supports 4K while producing images in 3–5 seconds, which suits high-volume work.
From the beginning it has won fans on one point: the images are simply beautiful. It is chosen by people who value artistic quality over technical accuracy.
V8 Alpha arrived on 17 March 2026, followed by V8.1 in April, which became the default in June. Rendering is five times faster, and 2K output is supported.
Released on 12 August 2026. Its features are that it can fix images as well as make them and that it can combine multiple reference images, which fits professional production workflows. The first reference image is free.
The FLUX.2 family arrived on 25 November 2025, taking over as flagship from FLUX.1. On 15 January 2026 a freely usable lightweight version, Klein 4B, followed. It is popular with people who want to run models on their own machines.
When it was released in 2022, it let anyone run image generation on their own computer, and that spread the field enormously. The culture of derivative models and fine-tuning also grew from here.
Stable Diffusion 4 arrived on 6 April 2026. It is the first generation since 2022 with a fundamentally new design, moving to the same approach as other companies' latest models.
The area with the fastest turnover this past year. It is normal for a model name from six months ago to be gone already. In 2026 the competition runs in three directions: making sound together with video, making long clips in one go, and making video from documents.
For quality, Seedance. To try things cheaply, Veo 3.1 Lite. To make video from documents, Wan 3.0. Prices in this area vary by more than 10 times, so start with something cheap to get a feel for it.
Version 2.0 arrived in February 2026 and took first place in video rankings, ahead of Google Veo 3, OpenAI Sora 2 and Runway Gen-4.5. Version 2.5 followed on 31 July.
With 2.5 you can make up to 30 seconds, including sound, in one pass. You can supply up to 30 reference images, 10 videos and 10 audio clips, and edit down to the second. The official output resolution is up to 720p, however. Resellers advertising 1080p or 4K may simply be upscaling it themselves, so be careful.
Released on 24 August 2026. Its main feature is that besides text and images, you can hand it documents, spreadsheets, slides, PDFs and web pages as they are, and it builds a video from them. No other company offers this idea of passing in your existing materials directly.
Clips run 2–30 seconds at 480p, 720p or 1080p. Failed generations are not charged, which is a real help in practice. Since the public beta on 6 August, it has been used for short dramas, ads, tourism promotion and music videos.
Veo was the first major model to generate dialogue, sound effects and ambient sound together with the picture. The current 3.1 comes in three quality levels, and the top and bottom differ in price by more than 10 times. If you do not need sound, Lite is among the cheapest video options of all.
On 30 June 2026 a new line, Gemini Omni Flash, also appeared. It handles text, images and video in one model, but its pay-as-you-go price has still not been published, which makes it hard to budget for work.
Released on 31 July 2026. It is designed to understand four kinds of input in a single model, and accepts up to 9 reference images, 3 videos and 3 audio clips. It produces 5–15 seconds at 2K and 24fps, with stereo sound.
In practice, resolution is 2K only. The price list does show a cheaper 768p tier, but it is invitation-only and cannot actually be selected through the API, so you are billed at the 2K price. Do not judge it as cheap from the price list alone.
Version 3.0 arrived on 4 February 2026. It can output 4K, 60fps, 15 seconds in one go and matches mouth movement in multiple languages. A faster Turbo version was added in mid-June. Pricing uses credits, and consumption depends on resolution and whether sound is included.
This company was in real video production back when AI video could only make a few seconds. Chinese companies are overtaking it on raw quality, but it keeps its fans thanks to fine-grained control for creators.
Known as Dream Machine. The price per second is nearly four times higher at 1080p than at 720p, so drafting at 720p and finishing at 1080p works well.
OpenAI's Sora 2 has ended. The app and website closed on 26 April 2026, and the developer API closes on 24 September. (This page previously said Sora ended on 24 March. That was the date it was announced. We have corrected it.) Once almost synonymous with AI video, it is no longer available. Its exit marked the shift of leadership in video toward ByteDance, Google and Alibaba.
Reading text aloud and creating voices themselves. Used for narration, audiobooks, dubbing and game dialogue. Since 2026, more of these tools do not just read aloud but act as someone you can talk to.
To have text read aloud, ElevenLabs. To have a conversation partner, GPT Live. These two serve completely different purposes, so start from what you want to do.
It became available in the ChatGPT app on 8 July 2026, and from 10 September 2026 you can also build it into your own service.
The key difference is that it can listen while it talks. Earlier voice AIs waited for you to finish before replying, and stopped if you talked over them. GPT Live listens even mid-sentence, so you can say "no, that's not what I meant" and the conversation keeps going. The response delay also dropped from 1.41 to 0.798 seconds.
In ChatGPT, Free plan users get a lighter version (mini), and Go, Plus and Pro users get the full version.
If you build it into your own service, read the pricing carefully. The listed 5 cents per minute covers only the listening-and-speaking part. The actual thinking is done by a separate AI, and that is billed separately. As a rough guide, a 10-minute conversation costs about $0.51 with a light model and about $0.86 with a smart one. On top of that, silent time is billed at the same rate.
One more thing: the 12 published voices are English accents plus Brazilian Portuguese and Filipino. No Japanese voice is listed. If you plan to use it in another language, test it before committing.
Started in 2022 by two childhood friends from Poland, it reached an $11 billion valuation in four years and raised $500 million in February 2026.
Its text-to-speech and voice cloning are especially well regarded and are widely used for audiobooks, dubbing, games and video production. There are speed-focused and quality-focused options, differing in price by about two times.
Two tiers: Turbo for speed and HD for quality. On its own it is priced about the same as ElevenLabs, but because MiniMax covers text, images, video and voice under one account, it becomes attractive if you want to try many things.
Music is a different story now. Since 20 August 2026, the music and lyrics generation APIs are not accepting new sign-ups. Existing users can continue, but new users must either use the MiniMax Audio app or run the published model on their own hardware.
Both are offered as part of the main platforms. Compared with the specialist ElevenLabs, they can fall a step behind on expressiveness, but if you already use them for text, you can start without a new contract.
On 26 August 2026 Google also released Gemini 3.5 Transcribe, specialized for transcription, as it builds out its voice lineup.
Then on 15 September came "Gemini 3.8 Live," a voice model you can hold a conversation with. Like GPT Live, you speak to it and it answers by voice. It is reported to support 97 languages, including Japanese, and it is priced low, at roughly 2.3 cents per minute for listening and speaking combined. The language information comes from third-party coverage and has not been confirmed in Google's own documentation, so try it before deciding.
Google's text-to-speech has four price levels depending on voice quality, with about a 40-times difference between top and bottom. The lowest level includes 4 million free characters per month, so small volumes cost nothing.
Creating whole songs and backing tracks. What makes this area different is that lawsuits over the rights to training data are still ongoing, so check the terms before using it for work.
For quality, Suno. For peace of mind in commercial use, ElevenMusic. ElevenMusic states that its training data is properly licensed.
As of February 2026 it had about 2 million paying users. Its output is highly rated across genres: pop, rock, electronic and ambient.
In this area, even the way services are offered is changing. MiniMax stopped taking new music API sign-ups on 20 August 2026 and moved to an app and an open model. If you plan to use one for work, check not only the price but also whether you will still be able to sign up in the future.
Meanwhile, lawsuits over training data are ongoing, and the conditions for commercial use are not settled. This is something to check and decide for yourself. There is no official API; access is through third-party providers.
ElevenLabs, which built its name in voice, moved into music in August 2025 and launched a standalone app on 1 April 2026.
The difference is less about quality and more about clear rights. It states that it trained on data 100% licensed from music rights organizations. It avoids the rights issues Suno and Udio face, so it is a strong option for commercial use.
Alongside Suno, one of the most highly rated services for output. It faces similar training-data lawsuits, so terms for commercial use are in flux. There is no pay-as-you-go API; only monthly subscriptions.
Google's music model. Besides the standard version that makes one song at a time, there is a real-time version that keeps the music flowing without a break, useful for things like continuous background music during a live stream.