AI

AI Roundup: Sora Officially Launches, Gemini 2.0 Debuts, With Agents and Productization in Focus

Updated · 2026-09-28 22:18 · 6 sources cited

AI Roundup: Sora Officially Launches, Gemini 2.0 Bets on Native Multimodality and Agents

OpenAI officially released Sora on the third day of its launch season. The model first appeared in February 2024 and launched after nearly 10 months of iteration, and is available to ChatGPT Plus and Pro subscribers.[1]

On top of text-to-video and image-to-video, Sora adds features such as storyboarding, adjusting original videos with text, and merging videos from different scenes. Plus users can generate up to 50 advanced videos at 720p and 5 seconds; Pro users can generate up to 500 advanced videos at 1080p and 20 seconds, and can remove watermarks.[1]

Google released Gemini 2.0 Flash, calling it the first model to achieve native multimodal input and output. DeepMind CEO Hassabis said it performs as well as 1.5 Pro, is twice as fast as 1.5 Pro, and surpasses 1.5 Pro on key benchmarks.[2]

Gemini 2.0 Flash can natively generate audio and images, supports multilingual TTS, image output, and conversational multi-turn editing, and can generate integrated responses of text, audio, and images through a single API call; all image and audio outputs use SynthID invisible watermarks.[2]

Google DeepMind scientist Nenad Tomasev said reinforcement learning frees AI from the limits of human knowledge; in the future, rather than relying on a single model, agents with multiple capabilities will be built. Kaggle CEO D.Sculley said AI progressed more in the past year than in the previous 7 years.[3]

On the startup side, AIGCode launched AutoCoder for product managers, which can generate software without requiring coding skills. The team evolved from MoE to a PLE hybrid multi-expert architecture, saying its 7B Xiyue large model is comparable to mainstream models such as GPT-4o in code.[4]

flomo co-founder Shaonan said he feels panicked about AI but is not in a hurry. After flomo added Related Notes and Find features, growth has been a relatively positive signal in recent years.[5]

LOOI is an AI hardware device unveiled by Chinese team TangibleFuture at CES, based on smartphone interaction. The team said that if viewed as a consumer product, it could score 70 points, but as what they truly want to build, it could only score 30; the first priority is to make it a successful consumer product.[6]

Analysis

From these developments, multimodal generation and agents have become the focus of competition among large models, while productization and deployment scenarios are receiving more attention.[1][2][3]

Sources

← All AI stories