Physical Intelligence π0.7, robot skill 재조합을 보였다
Physical Intelligence는 π0.7이 task별 specialist training 없이도 새 language command와 unseen task를 처리하는 초기 compositional generalization을 보였다고 밝혔다. Laundry folding에서는 UR5e task data 없이 expert teleoperator의 zero-shot success와 맞먹었다.
원문: π 0.7: a Steerable Model with Emergent Capabilities 원문 보기 →
Physical Intelligence의 π0.7은 robot demos가 자주 막히던 병목을 겨냥한다. task마다 별도 specialist model을 만들고 data를 다시 모으는 방식에서 벗어날 수 있는지다. 회사는 2026년 4월 16일 research post에서 π0.7을 general-purpose vision-language-action model로 설명하며, training data에 없던 새 language command와 task를 수행할 수 있다고 밝혔다.
핵심은 compositional generalization이다. Physical Intelligence는 π0.7이 여러 task에서 배운 skill을 recombine해 new kitchen appliances 사용 같은 문제를 풀고, laundry folding data가 없는 new robot에서도 folding을 수행했다고 설명했다. 회사는 이를 LLM이 알고 있는 개념을 새 형식으로 조합하는 능력에 비유하지만, robotics에서는 physical motion, robot morphology, scene variation이 들어가 훨씬 더 다루기 어렵다.
가장 중요한 detail은 UR5e transfer
π0.7은 bimanual UR5e system에서 laundry를 fold하도록 평가됐다. Source robot과 UR5e는 size, positioning, morphology가 크게 다르고, 회사는 이 task에 대한 UR5e training data를 모으지 않았다고 적었다. 그럼에도 π0.7의 success rate는 source robot에서 data를 수집했던 expert human teleoperators가 UR5e에서 처음 시도했을 때의 zero-shot success rate와 맞먹었다. 그 teleoperators의 평균 teleoperation experience는 375 hours였다.
방법론도 단순히 더 큰 dataset만을 말하지 않는다. π0.7은 language, metadata, control modality labels, visual subgoal images처럼 다양한 prompt structures를 training에 넣는다. 이 prompt는 무엇을 할지뿐 아니라 어떻게 할지를 지정한다. Test time에는 standard language instructions 외에도 desired strategy와 lightweight world model이 만든 visual subgoal을 받을 수 있다.
다만 이 결과를 deployed robot product로 읽으면 안 된다. Source는 “first signs”와 “initial signs”라는 조심스러운 표현을 쓴다. 아직 외부 replication, standardized robotics benchmark, 비용과 failure mode 공개가 필요하다. 그래도 의미는 크다. 만약 한 model이 task-specific specialist와 비슷한 성능을 내면서 unseen combinations를 다룰 수 있다면, embodied AI의 bottleneck은 task마다 새 model을 만드는 일에서 instruction design, safety envelope, evaluation으로 이동할 수 있다.
관련 기사
중국 휴머노이드 로봇, VLA로 대규모 공장 수주 — 상반기 출하량 272% 급증
UBTECH, Galbot, AgiBot 등 중국 업체들이 BYD·Foxconn·CATL에 수백억 원 규모 수주를 달성했다. VLA 아키텍처 덕분에 별도 재학습 없이 언어 지시만으로 작업 전환이 가능한 점이 핵심 경쟁력으로 꼽혔다.
PhyFilter, 적은 데이터로 낯선 지형·바람 적응…로봇 4종서 물리 피드백 검증
학습 정책의 출력을 물리 기반 피드백으로 바로잡는 PhyFilter가 사족보행 로봇·드론·공중 매니퓰레이터 등 4종 시스템에서 낯선 조건 적응을 보였다. 대규모 시연 데이터를 계속 늘리는 방식 대신 가벼운 보정 모듈로 일반화를 높였다는 결과다.
Reddit를 달군 Generalist GEN-1, 단순 robot task 99% success 주장
Generalist는 GEN-1이 더 높은 success rate, 빠른 execution, 낮은 task-specific robot data 요구량을 통해 단순 physical task의 commercial threshold를 넘기 시작했다고 말한다.