AI

From Personal AI Assistant to Brain Science Infra: Agents Begin to Fill In Touch, Memory, and the Physical World

Updated · 2026-10-10 10:05 · 6 sources cited

From Personal AI Assistant to Brain Science Infra: Agents Begin to Fill In Touch, Memory, and the Physical World

Recently, progress in AI is no longer just a contest over model parameters and leaderboard scores; it is simultaneously penetrating four directions: personal daily life, brain science infrastructure, embodied manipulation, and enterprise deployment. From a ten-day experience with Today AI, to Kunwei's ultrasound-through-skull brain reading, to new moves by Sharpa, Robot Era, MirroS, and openJiuwen, a common thread is becoming clear: for agents to truly enter the real world, they must fill in the system capabilities of perception, memory, action, and safe operation.

Personal AI assistants first learn to be 'proactive' and stay in 'the same conversation'

Today AI was created by Qi Junyuan, founder of Teambition. Its domestic version launched on September 24 and has connected to commonly used Chinese apps such as Feishu, DingTalk, Tencent Docs, and email [1]. After a ten-day trial, the author found that what makes it Personal is not just answering questions, but proactively generating morning and evening briefings based on personal information provided by the user, pushing updates about favorite teams, email summaries, and cover articles from publications; in the same chat box, a user can gather topic research materials one second and ask whether a certain game has been released the next, and Today can pick up the thread and return to work tasks [1]. The desktop version puts conversations, briefings, tasks, memory, and linked apps in the same interface. After connecting Feishu, it can read work documents from the past week and push topic ideas every morning at 9 [1]. In a heavier task, the author sent Today a zip package for a game mod; it unpacked the package itself, read the documents, and called Claude Opus 5 to begin localization, but for the final step, 'where to put the localization files in the game directory', it repeatedly gave confident but ineffective instructions [1]. This shows both the possibility of Personal AI using long-term memory and proactive service to lower interaction costs, and its reliability shortcomings in a real operational loop.

Using low-frequency ultrasound to 'read the brain', Kunwei aims to be the infra for brain science AGI

Shenzhen-based Kunwei Technology recently completed a new round of financing worth 400 million yuan, with new shareholders including XVC and Sequoia China, while existing shareholders Sherpa and Tencent further increased their positions [2]. Its technological path abandons the industry habit of constantly increasing frequency to gain clarity, turning instead to low frequency, bringing computing power and algorithms into the physical acoustic field, and balancing skull penetration with high-resolution real-time imaging in the low-frequency band [2]. In tests, top overseas flagship color ultrasound machines could only piece together a blurry outline after penetrating the skull, while Kunwei's device could clearly present the midbrain nuclei, lateral ventricles, and even blood flow dynamics in tiny blood vessels [2]. Founder Liu Jiajia summarizes this path as 'acoustic spatial intelligence' and says it replicates bats' acoustic spatial perception: the raw echo data stream generated by ultrasound reaches as high as 10GB per second, and the model must parse voxel reflection features in three-dimensional space in real time with an extremely small parameter count [2]. Kunwei has already achieved tens of millions in orders, and its goal is not limited to transcranial ultrasound equipment, but to provide hardware and the foundational software platform kOS, becoming an infra company that empowers brain science AGI [2].

Touch becomes the gateway to dexterous manipulation, Sharpa releases three products at once

At IROS, Sharpa released its first fully self-developed general-purpose humanoid robot D01, along with a new generation of fully tactile, ultra-compact, lightweight dexterous hand W02 and a high-fidelity exoskeleton tactile data glove AE01 [3]. Several numbers for D01 revolve around dexterous manipulation: the arm payload-to-self-weight ratio is close to 1:1, maximum end-effector speed exceeds 10.5m/s, communication frequency is 1000Hz, repeat positioning accuracy is 0.2mm, and the spherical wrist joint uses a 1:1 human-scale design; electronic skin covers the whole body, tactile sensing covers the upper body, tactile sampling rate is 100Hz, force sensing range is 0.1—20N, and force resolution is 0.2N [3]. W02 instead reduces active degrees of freedom from 22 in the previous generation W01 to 21, shrinks the whole-hand size by about 30%, and achieves 'zero grasping radius'; its fingertip visual-tactile sensor has a force sensing range of 5mN—30N and spatial resolution of 1mm [3]. AE01 uses 22 encoders to capture the operator's natural hand movements and feeds robot-side tactile status back to the person, for precise teleoperation and first-person-view data collection [3]. These products together point to one judgment: after a robot enters a real workspace, vision can only tell it that 'there is something ahead', while touch can answer 'where it was touched, how much force is applied, and whether there is slipping'.

World action model tops the leaderboard: Robot Era VPP2 separates prediction and action

Robot Era's self-developed world action model VPP2 recently topped the RoboDojo simulation leaderboard, with an overall average success rate of 32.26% and an overall average score of 39.26 points, ranking first in both metrics; VPP2 also took first place in the three dimensions of generalization ability, fine manipulation, and memory ability [4]. By comparison, in post-training evaluation on the RoboDojo simulation benchmark, GPT-6-Astra had an average success rate of 22.48% and an average score of 28.97 points [4]. VPP2 added no extra data and used no enhancement methods such as Agent RSI, improving only with standard datasets, indicating that its performance comes from pretraining and base-model generalization [4]. On the technical path, the team trained video prediction and action learning in stages. It first used Alibaba's open-source Wan2.1-I2V-14B as a foundation, integrating data such as robot manipulation, human activities, and general video, so the model learns the changes in a complete operation from start to finish; in instruction-following tests on robot manipulation videos, the 14B-parameter VPP2 reached a 90% success rate, while the 64B-parameter Cosmos3 reached 78% [4]. After that, VPP2 introduced a 0.9B-parameter diffusion Transformer as an action expert. In the early stage of action training, it froze the video model's base parameters and adapted only through LoRA to preserve generalization ability; in the end, video prediction took about 0.12 seconds, the action expert took about 0.1 seconds, and the overall action clip generation latency was about 0.22 seconds [4]. On a real ALOHA dual-arm robot, VPP2 achieved an average success rate of 58.5% across 10 categories of zero-shot tasks including grasping, placing, stacking, folding, and pouring, higher than π0.5's 40% [4].

Code builds the world, diffusion paints reality: AgentGarten lets agents evolve while playing

The MirroS team released AgentGarten, connecting an executable code environment with a real-time neural renderer: code defines physical rules, the model generates visual feedback, and throughput above 30 fps lets agents continuously interact with the world [5]. In the classic hide-and-seek experiment, hiders learned by round 4 to move barriers to build shelters, and seekers learned by round 10 to use ramps to climb over walls; after a failed jump, the seeker would also proactively push the ramp closer, adjust its position, and try again [5]. As a comparison, OpenAI's 2019 research relied on reinforcement learning from scratch and took about 25 million episodes to figure out shelter building, and about 100 million episodes to learn to use ramps to climb over walls [5]. AgentGarten's solution is to give physics to code and light and shadow to the neural model: after an agent gives an action command, the code environment immediately calculates collisions and state updates, and exports a lightweight geometric sketch from a first-person perspective; the neural renderer then combines visual memory to generate the next frame in real time [5]. The team also proposed the Adversarial Forcing training method and an exact replay mechanism to address error accumulation in streaming generation and temporal consistency challenges in long sequences [5]. This 'explore—review—consolidate' loop shifts the cost of expanding virtual environments from art labor-intensive to code-procedural expansion.

Enterprise-grade AgentOS goes open source: from usable to manageable, controllable, and scalable

openJiuwen released and open-sourced an enterprise-grade AgentOS, built by teams including Huawei's 2012 Labs, Huawei Cloud, Computing, Devices, and the Computing Power Vanguard Team, together with universities and enterprise developers. It integrates multi-agent collaboration, self-evolution, computing power affinity, and enterprise-grade security and reliability into a unified foundation [6]. Through swarm collaboration and workflow orchestration, it supports task decomposition, role division, and parallel execution, and allows humans to participate in judgment at key steps; relying on execution trajectories, user feedback, and memory mechanisms, agents can optimize skills and collaboration methods and consolidate effective experience into reusable capabilities [6]. At the runtime level, AgentOS improves computing power utilization efficiency through software-hardware collaboration, resource scheduling, and inference optimization, and constrains data access and tool calls with multi-tenant isolation, secure sandboxes, permission management, and safety guardrails, while supporting highly reliable continuous task operation with mechanisms such as fault recovery [6]. At HC2026, Huawei released an all-in-one open-source Agent acceleration platform and solution based on Jiuwen AgentOS. The platform includes a unified Agent gateway, a distributed runtime foundation, and the WorkSwarm unified workbench for office/programming, enabling hour-level end-to-end private deployment for enterprises [6]. The plug-in architecture and open northbound ecosystem allow five types of assets—skills, connectors, plug-ins, experts, and expert teams—to be flexibly combined, and published, acquired, and shared through Agentic Hub [6].

Common trend: agents are filling in the 'body' and the 'foundation'

Putting these threads together, personal AI assistants are competing for the daily entry point of users, brain science ultrasound and embodied intelligence are competing for perception and manipulation capabilities in the physical world, and enterprise-grade AgentOS is solving multi-agent collaboration, security, and deployment problems. Their common point is that AI is no longer just a model passively answering questions, but is trying to remember goals over the long term, proactively acquire information, act through touch and vision loops, and be manageable, isolatable, and recoverable in enterprise environments. What is worth watching next: whether personal assistants like Today AI can push 'light and fast conversation' into reliable heavy-task execution; whether Kunwei's kOS can be truly invoked by brain-computer interface labs and NeuroAI model companies; whether Sharpa's tactile data, Robot Era's VPP2, and AgentGarten's real-time environment can move robots from simulation leaderboards and demos to stable commercial value; and whether openJiuwen's AgentOS can turn hour-level private deployment into scaled implementation of industry agents. [1][2][3][4][5][6]

Sources

← All AI stories