LocalLLaMA 화제: Gemma 4 31B의 FoodTruck Bench 약진을 둘러싼 논쟁
LocalLLaMA 스레드가 Gemma 4 31B의 예상 밖 FoodTruck Bench 성과를 끌어올렸다. 토론은 곧 장기 계획 능력과 benchmark 신뢰성 문제로 이어졌다.
태그
#llm 태그가 달린 기사
LocalLLaMA 스레드가 Gemma 4 31B의 예상 밖 FoodTruck Bench 성과를 끌어올렸다. 토론은 곧 장기 계획 능력과 benchmark 신뢰성 문제로 이어졌다.
Anthropic의 새 interpretability 연구는 Claude Sonnet 4.5 내부의 감정 관련 표현이 특히 스트레스 상황에서 행동을 바꾸는 인과적 역할을 한다고 주장한다.
Hacker News에서 주목받은 새 논문은 verifier나 teacher model, reinforcement learning 없이도 모델이 자기 답안을 바탕으로 코드 생성 성능을 높일 수 있다고 주장한다. 논문은 Qwen3-30B-Instruct가 LiveCodeBench v6 pass@1에서 42.4%에서 55.3%로 상승했다고 보고했다.
Stanford의 공개 CS25 강의는 Zoom, recordings, Discord를 통해 campus 밖까지 확장된 Transformer 연구 학습 채널로 다시 작동하고 있다.
Lemonade는 GPU·NPU를 겨냥한 OpenAI-compatible server로 local AI inference를 패키징해, everyday PC에서 open model 배포를 더 쉽게 하려는 스택이다.
r/LocalLLaMA에서 CoPaw-9B 관련 글이 142점과 29개 댓글을 기록하며 주목을 받았다. 스레드는 Qwen3.5 기반의 9B Agent 모델, 262,144 token context, 그리고 GGUF·quantized 배포 가능성에 대한 관심을 중심으로 반응이 갈렸다.
Cloudflare는 2026년 3월 30일 advanced Client-Side Security 도구를 전체 사용자에게 개방했다고 밝혔다. Cloudflare 블로그에 따르면 이번 업데이트는 graph neural network와 LLM triage를 결합해 false positive를 최대 200배 줄이고, advanced 기능을 self-serve로 열면서 free bundle에 domain-based threat intelligence도 포함한다.
Mistral이 2026년 3월 16일 Mistral Small 4를 공개했다. 119B total parameters, 6B active parameters, 256k context window, Apache 2.0, configurable reasoning_effort를 결합해 reasoning·coding·multimodal 작업을 한 모델에 모았다.
Mistral이 2026년 3월 16일 Lean 4 전용 오픈소스 코드 에이전트 Leanstral을 공개했다. 6B active parameters, Apache 2.0 공개, FLTEval 도입, Mistral Vibe와 API 및 가중치 배포가 핵심이다.
r/MachineLearning의 새 글이 TurboQuant를 KV cache 논의에서 weight compression 단계로 끌어왔다. GitHub 구현은 low-bit LLM inference용 drop-in path를 목표로 한다.
Google Research는 2026년 3월 16일 high-temperature superconductivity 질문 67개로 여섯 개 LLM 시스템을 평가한 결과를 공개했다. NotebookLM과 custom RAG처럼 curated reference를 쓰는 폐쇄형 구성이 open-web 모델보다 더 높은 점수를 받았다.
Google는 Mar 03, 2026, Gemini 3.1 Flash-Lite를 Gemini 3 series 중 가장 빠르고 cost-efficient한 model로 공개했다. preview 단계부터 낮은 token 가격과 높은 throughput을 앞세워 대량 developer workload를 겨냥한다.