In partnership with

AI help, without the trust tax.

Most AI tools ask you to trade your data for intelligence. Norton Neo doesn't. It's the first safe AI-native browser built by Norton, and it gives you powerful built-in AI without handing your privacy over to get it. Search, summarize, and write with AI built directly into your browser. Your data stays yours. Your context stays private.

Built-in VPN, anti-fingerprinting, and ad blocking come standard. No add-ons. No setup. No compromises.

Fast. Safe. Intelligent. That's Neo.

Featured

Black Forest Labs Just Built a Model That Learns the World, Not Just Its Projections.

FLUX 3 jointly trains on images, video, and audio at the same time — because no single modality captures reality completely. The result already beats every major video generator in head-to-head comparisons. The robotics application might be the bigger story.

There's a sentence near the top of Black Forest Labs' FLUX 3 announcement that does something unusual: it explains, in plain language, a genuinely important idea about how intelligence works. "Learn from one and you get a good model of that projection. Learn from all of them at once and their mutual constraints tell you more." The projection framing is precise. An image is a projection of reality — it captures spatial relationships at one moment but loses time. A video restores time but loses the sounds that reveal causality. Audio reveals what vision can't: the crack that tells you something broke, the footstep that tells you something arrived. Language links all of it to goals and abstractions. No single modality is reality. Each is a partial view of it.

FLUX 3 — launched July 23rd in Early Access by the team that built FLUX 1 and FLUX 2, the open-weight image generators that dominated the independent creative AI ecosystem — is the first model the company has built on the principle that you should train on all of them simultaneously, using a unified architecture called Self-Flow, so the modalities teach each other about the underlying reality they're each partially encoding.

WHAT'S HAPPENING: One model, images, video, synchronized audio, and robot action — all from the same backbone

FLUX 3 is capable of generating images, videos up to 20 seconds with synchronized native audio, and action predictions — from text prompts, image references, video references, or combinations. The video capabilities span text-to-video, image-to-video, video-to-video character transfer, generative video-audio continuation, keyframe-to-video, multilingual dialogue, a broad range of visual styles, and agentic chaining of individual clips into longer multi-shot sequences. The benchmark numbers are preliminary but striking: FLUX 3 was preferred over Runway Gen-4.5 in 77% of head-to-head comparisons, over Luma Ray 3.2 in 93% of comparisons, over Grok Imagine Video in up to 69%, Kling v3 Pro in 60%, and Happy Horse v1 in 59%. FLUX 3 is particularly strong in capturing human facial expressions, associating sounds with physical events, and multilingual generation — capabilities that require the model to have learned something real about how faces and mouths and voices relate to each other across modalities.

WHY IT MATTERS: The same backbone that generates video is being used to control robots at Audi

The most forward-looking section of the announcement isn't the video benchmark. It's the action prediction application. Black Forest Labs partnered with mimic robotics to develop FLUX-mimic, a video-action model that combines the FLUX 3 backbone with mimic's expertise in robot learning for dexterous manipulation — and it's being tested on real production tasks at Audi. The thesis Black Forest Labs is advancing here is architectural: the same model that learns to predict what a video will look like next — because it understands physical causality, object persistence, and how the world evolves through time — is also a foundation for learning how to act in that world. Physical AI and content creation run on the same underlying representation problem. The world model that makes video generation coherent is the world model that makes robot manipulation plannable.

"The modalities stop being separate and start being evidence about one underlying reality." — Black Forest Labs, FLUX 3 announcement

THE BIGGER PICTURE: Black Forest Labs is building toward a unified perception-action-language model

The FLUX 3 announcement is explicit about where this is heading: the goal is to unify perceptual, action, and language prediction in the same unified model — the next step beyond a model that understands images, video, and audio. That's a direct statement of intent to build something closer to general visual intelligence than any image or video generator the company has shipped before. The launch plan reflects a phased approach: video and audio generation in Early Access now, image synthesis in the following weeks, open-weight access via FLUX 3 Dev to follow. Black Forest Labs built FLUX 1 as an open-weight model that became the foundation for an enormous creative ecosystem — hundreds of LoRA fine-tunes, community tools, ComfyUI integrations. FLUX 3 Dev would bring multimodal video-audio-image generation into that same open ecosystem, which would be a significant moment for independent developers.

MY TAKE: This is what "world model" means when it's built from the ground up, not bolted on

The phrase "world model" has been used in AI research to mean many things — from the simple "internal representation that helps predict consequences" to the grand "a model that understands physical reality." Black Forest Labs' architecture paper makes a specific and falsifiable claim: that jointly training on images, video, and audio produces better representations than training on each separately, because the cross-modal constraints function as a form of self-supervision. The sound has to match the impact. The motion has to obey the mass. The future has to follow from the past. If the model learns to satisfy those constraints, it's not just memorizing patterns — it's learning something about how the world works.

The robotics application is the proof-of-concept that makes the theoretical argument concrete. A model that generates video coherently must have learned something about how objects behave in physical environments — how they collide, how they move, what they look like from different angles. Those are exactly the representations a manipulation robot needs to plan actions. Mimic robotics getting early access to FLUX 3 and applying the backbone to dexterous manipulation tasks at Audi is real-world validation of the underlying thesis, not a demo.

The Starchild-1 story from a few weeks ago covered Odyssey doing something adjacent — a world model that generates synchronized audio and video in real-time. The architectural approach is different, but the underlying bet is similar: that the next generation of AI systems will ground their intelligence in a richer, multi-sensory representation of the world, rather than a text-centric or single-modality one. Two companies arriving at the same insight independently, in the same month, is a stronger signal than either one would be alone.

So here's the question worth sitting with: if the same model architecture that learns to generate coherent video — by learning the physical constraints that connect sound to motion, motion to cause, cause to effect — is also the right foundation for robots that need to act in physical environments, what does that tell us about what intelligence actually is?

Source: Black Forest Labs — "FLUX 3: Real World Models," July 23, 2026

PRDs by voice. Bug reports by voice. Ship faster.

Dictate acceptance criteria and reproductions inside Cursor or Warp. Wispr Flow auto-tags file names, preserves syntax, and gives you paste-ready text in seconds. 4x faster than typing.

See the whole platform. No guided tour.

Skip the sales call. Walk through Gladly's interface yourself — the AI suggestions, the unified customer view, the full conversation thread. 15 minutes, no installation, no commitment.

  • Sora - officially launches to the public - create videos from prompts or images

  • Claude - Tackle any big, bold, bewildering challenge with Claude

  • Fireflies.ai - AI notetaker and transcription for meetings!

  • Taskade - Create and Train your own AI Agents!

  • AI Tools for Bloggers - Leveraging AI Tools and Pinterest for Success

  • ChatGPT - What will it do for you?!

  • Grok - Harness powerful AI & generate stunning images

  • Gemini 2.0 - Faster and more capable than ever!

  • Replit - Take your ideas and turn them into software — no coding required!

  • Submagic - lets you create viral shorts in seconds!

  • Midjourney - create incredible images from basic prompts!

  • MadeByMelo - An inclusive & collaborative space for artists, creators, & gamers

You got a minute?

You got a minute?

Your cozy spot to learn how to focus better, work smarter, and take care of yourself - all things AI, productivity, & mental wellness.

The Rundown AI

The Rundown AI

Get the latest AI news and learn how to use it to get ahead in your work and life. Join 2,000,000+ readers from companies like Apple, OpenAI, and NASA.