Perplexity, Qwen SFT+RL로 GPT factuality 비용 곡선 추월 주장
중요한 점은 검색형 AI가 유창한 답변보다 factuality와 citation 품질로 평가된다는 데 있다. Perplexity는 SFT + RL pipeline으로 Qwen model이 더 낮은 비용에서 GPT model의 factuality를 맞추거나 앞선다고 주장했다.
원문: Perplexity said SFT and RL post-training let Qwen models match or beat GPT factuality at lower cost 원문 보기 →
tweet가 드러낸 점
Perplexity는 최신 model 작업을 chat style이 아니라 search quality로 설명했다. 핵심 quote는 Our SFT + RL pipeline improves search, citation quality, instruction following, and efficiency. With Qwen models, we match or beat GPT models on factuality at a lower cost. 이다.
Perplexity account는 AI search product release, app update, research note를 주로 올리는 공식 채널이다. 이 tweet가 material한 이유는 training recipe와 evaluation target, 비교 대상을 함께 적었기 때문이다. supervised fine-tuning과 reinforcement learning을 거친 Qwen model이 factuality와 cost 측면에서 GPT model과 경쟁한다는 주장이다.
왜 의미가 있나
search-augmented assistant의 실패는 일반 chat benchmark에서 잘 드러나지 않는다. 답변은 매끄럽지만 source가 약하거나, 새 문서를 놓치거나, 싼 query에도 비싼 model을 쓰는 문제가 생길 수 있다. Perplexity의 claim은 search behavior, citation quality, instruction following, efficiency라는 production 변수 네 가지를 동시에 겨냥한다.
FxTwitter metadata 기준으로 이 tweet에는 public paper, repo, blog URL이 붙어 있지 않고 media attachment만 확인된다. 따라서 결과는 Perplexity가 보고한 benchmark로 취급해야 하며, 독립 검증 전에는 method를 단정하기 어렵다. 그래도 signal은 분명하다. Qwen 계열 open model이 단순히 저렴한 inference backend가 아니라, closed GPT-class system과 factuality layer에서 경쟁할 수 있는 trainable search model로 포지셔닝되고 있다.
builder 관점의 다음 질문은 method다. 어떤 factuality dataset을 썼는지, citation은 human review인지 automatic check인지, 개선분이 retrieval policy에서 온 것인지 answer model fine-tuning에서 온 것인지가 중요하다. cost도 per query, per token, successful answer, latency target 중 무엇을 기준으로 했는지 확인해야 한다. 다음 관전점은 Perplexity의 technical write-up, model card, 혹은 실제 traffic routing 변경이다.
Source: X source tweet
관련 기사
Qwen3.8-Max-0902, CodeArena 1,691점으로 1위…100만 토큰 유지
Qwen3.8-Max의 9월 2일 체크포인트가 프런트엔드 CodeArena에서 1,691점으로 선두에 올랐다. 기반 규모를 바꾼 신모델이 아니라 코딩과 장기 에이전트 작업을 겨냥한 후훈련 갱신이며, 100만 토큰 문맥과 API 호환성을 그대로 유지한다.
Qwen3.6 GGUF 논쟁, r/LocalLLaMA는 “어떤 quant를 돌릴 것인가”로 내려갔다
r/LocalLLaMA가 Qwen3.6 release 자체보다 GGUF quant 선택과 CUDA 버그에 더 크게 반응했다. Unsloth의 benchmark post는 KLD, disk space, 4bit gibberish, CUDA 13.1/13.3 같은 실제 실행 조건을 전면에 올렸다.
LocalLLaMA 화제: 듀얼 RTX PRO 6000 Blackwell에서 Qwen3.5-122B 198 tok/s 검증
LocalLLaMA에서 주목받은 글은 SGLang b12x+NEXTN, PCIe switch topology, 공개 raw benchmark JSON을 바탕으로 듀얼 RTX PRO 6000 Blackwell에서 Qwen3.5-122B NVFP4가 약 198 tok/s를 기록했다고 공유했다.