본문으로 건너뛰기

카테고리

LLM

Insights의 LLM 기사

RSS 피드
LLM 레딧

r/LocalLLaMA가 본 NVIDIA Nemotron 3 Super 공개

NVIDIA의 Nemotron 3 Super는 120B total / 12B active hybrid Mamba-Transformer MoE, native 1M-token context, 그리고 open weights·datasets·recipes를 함께 내세운다. LocalLLaMA discussion은 이 openness와 efficiency claim이 실제 home-lab deployment로 이어질 수 있는지에 집중했다.

1분 소요 38 조회
LLM 레딧

r/LocalLLaMA가 주목한 llama.cpp reasoning budget 제어

새로운 llama.cpp 변경은 <code>--reasoning-budget</code>를 template stub이 아니라 sampler 차원의 실제 제어로 바꾼다. LocalLLaMA thread는 긴 think loop를 줄이는 것과 answer quality를 지키는 것 사이의 tradeoff, 특히 local Qwen 3.5 환경에서의 의미를 집중적으로 논의했다.

1분 소요 52 조회
LLM X/Twitter

OpenAI, Responses API용 컴퓨터 환경 설계 원칙 공개

OpenAI Developers는 2026년 3월 11일 글에서 Responses API가 장시간 agent workflow를 처리하기 위해 hosted computer environment를 어떻게 구성했는지 설명했다. 핵심은 shell execution, hosted container, 통제된 network access, reusable skills, 그리고 native compaction이다.

2분 소요 43 조회
LLM X/Twitter

NVIDIA, multi-agent AI용 Nemotron 3 Super 공개

NVIDIA AI Developer는 2026년 3월 11일 Nemotron 3 Super를 공개하며, 12B active parameters를 사용하는 오픈 120B-parameter hybrid MoE 모델과 native 1M-token context를 강조했다. NVIDIA는 이 모델이 이전 Nemotron Super 대비 최대 5배 높은 throughput으로 agentic workload를 겨냥한다고 설명했다.

2분 소요 49 조회