Back to Blog

OpenAI Ships GPT-6 Sol and Luna, Google Adds TTS and Call Agent

OpenAI cuts API costs with two new models, Google launches prompt-designed voices and a Gemini phone-call agent, and Meta's Muse leaks its own filesystem

OpenAI Ships GPT-6 Sol and Luna, Google Adds TTS and Call Agent
This Week's Key Stories
  • OpenAI released GPT-6 Sol at $2/$10 and Luna at $0.10/$0.50 per million tokens, both available now in the API.
  • Google launched Gemini 3.8 Flash TTS with prompt-based voice design across 100-plus languages, ranking first on Hume AI's Voice Design Benchmark.
  • Google Gemini can now make phone calls to local businesses on your behalf directly from Pixel 11, no hold music required.
  • NVIDIA released Nemotron 3 Diarization, a 100M-parameter open-weight model that tracks up to 8 overlapping speakers in real time.
  • Meta's Muse was found to share its entire root filesystem with users after minimal prompting, exposing Ubuntu system files and app templates.
  • Midjourney added live style previews and edit updates to its alpha interface, letting you see your prompt across styles before committing.
01

OpenAI Ships GPT-6 Sol and Luna at Lower Cost

OpenAI released two new models this week: GPT-6 Sol and GPT-6 Luna. Sol is priced at $2 input and $10 output per million tokens. Luna sits at $0.10 input and $0.50 output, making it one of the cheapest capable models in the API right now. Both are trained with methods similar to GPT-6 Astra and are available immediately in the API, ChatGPT Work, and Codex.

For creators building agents or automations, Luna's pricing means you can run far more calls for the same budget. Sol sits in a practical middle tier for tasks that need more reasoning than a lite model but do not justify Astra costs. Both models come with improved prompt caching, which matters a lot for long-running agent loops where context gets resent repeatedly.

If you have been holding off on building API-based tools because costs felt unpredictable, this is the week to revisit that. Full pricing and benchmark details are here.

$0.10
Luna input cost per 1M tokens
50%
Cheaper than previous tier
02

Google's New TTS Designs Voices From Plain Text

Google released Gemini 3.8 Flash TTS and Flash-Lite TTS this week through the Gemini API and Google AI Studio. The headline feature is prompt-based voice design: you describe the voice you want in plain language and the model generates it. No cloning, no recording, no preset selection. It works across more than 100 languages and scored 71.4 on Hume AI's Voice Design Benchmark, putting it at the top of that leaderboard.

Flash-Lite is the cheaper, faster sibling for high-volume use cases where cost per character matters more than absolute quality. Flash is the one to reach for when you need a voice that actually sounds like a deliberate creative choice rather than a default. For anyone producing podcasts, narration, or voiceovers at scale, the ability to iterate on voice character through prompts rather than recording sessions is a real workflow shift.

Both models are available now. Read the full breakdown of capabilities and API access here.

TRY IT ON SNIPPT

Try Snippt's text-to-speech tool to turn your script into a finished voiceover without any recording setup. Text to Speech →

71.4
Score on Hume AI Voice Design Benchmark (ranked #1)
100+
Languages supported
03

Gemini Will Make Your Phone Calls for You

Google is rolling out an early experiment on Pixel 11 that lets Gemini call local businesses on your behalf. You tell it what you need, like a restaurant reservation or a stock check, and Gemini makes the call and handles the conversation while you do something else. You do not even need to start the call yourself.

This is the most practical version of the AI agent pitch that has shipped to consumers so far. Waiting on hold is a real cost in time, and delegating it to a model that can handle natural conversation is a straightforward win. The limitation right now is that it is Pixel 11 only and framed as an experiment, so broad availability is not guaranteed.

For creators who also run small businesses or freelance operations, this kind of task delegation is worth watching closely as it expands. The Verge has the details on how it works.

04

NVIDIA Drops Open-Weight Speaker Diarization Model

NVIDIA released Nemotron 3 Diarization on Hugging Face this week. The 100M-parameter open-weight model answers one question about any audio: who spoke when. It tracks up to 8 speakers, handles overlapping voices, and runs on both pre-recorded files and real-time streams from a single checkpoint.

For creators doing interview content, podcast production, or any multi-speaker transcription workflow, accurate diarization is the piece that usually breaks first. Getting speaker labels wrong means manual cleanup that eats time. A model this small that handles overlap and real-time simultaneously is genuinely useful at the local or API level without the cost of a larger system.

The checkpoint is on Hugging Face now. Full technical details and model access are here.

100M
Parameters
8
Speakers tracked simultaneously
05

Meta's Muse Hands Over Its Own Filesystem

Meta's Muse AI agent had a rough week on the security front. Developers Peter James and Jonny L. Saunders independently found that with minimal prompting, Muse would zip up and share the entire contents of its root filesystem, including Ubuntu system files and app templates. The two researchers say they found this without any sophisticated jailbreaking.

This matters for creators because Muse had topped the App Store charts with an estimated 600,000 daily active users in the US. A lot of people are trusting it with their workflows. A consumer-facing agent that exposes its own infrastructure to casual prompting is a signal to be careful about what you share with any agent tool, not just this one.

Separately, The Verge reported the full details of the filesystem exposure. Meta has not yet commented publicly on a fix timeline.

600,000
Estimated daily active US users at time of disclosure
06

Midjourney Adds Live Style Previews to Alpha

Midjourney pushed a set of updates to its alpha interface at alpha.midjourney.com this week. The most useful addition for active users is live style previews in the Styles sidebar: turn on the feature and you see a thumbnail of your current prompt rendered across different styles before you commit to a generation. That cuts the trial-and-error loop that burns credits when you are hunting for the right aesthetic.

The update also includes edit improvements, though Midjourney kept the specifics brief in its release notes. Fast model experiments are also running in the interface. The full update post is here. If you have been on the alpha, the live previews are worth turning on immediately.

TRY IT ON SNIPPT

Use Snippt's image generator to test prompt variations quickly alongside your Midjourney explorations. Image Generator →

07

What This Week Means for Creators

The pricing moves from OpenAI and the voice tools from Google are the practical story this week. Luna at $0.10 per million input tokens removes the main excuse for not building API-powered tools, and Google's prompt-based TTS means voice character is now a writing problem, not a recording problem. Both of those lower the floor on what you need to get started with a real production workflow.

The Muse filesystem story is a reminder that consumer AI agents are still early infrastructure running in public. Be deliberate about what you feed into any agent tool right now, and assume that anything you type could surface in ways the product team did not plan for. The models are getting cheaper and more capable fast. The security layer is catching up slower.

Try These AI Tools Today

From AI video to voice generation, Snippt gives you access to the latest models in one place.

Explore All Tools