The  LGTM
  • Home
  • Agentic Coding
  • Claude Code
  • Codex
Sign in Subscribe

Qwen3-8B

A collection of 2 posts
DeepRubric Shows Deep Research Agents Need Better Rewards Before They Need More Rollouts
ai-models

DeepRubric Shows Deep Research Agents Need Better Rewards Before They Need More Rollouts

Deep research agents do not mainly need longer traces. They need rewards that can tell the difference between grounded synthesis and beautifully cited filler. That is why DeepRubric is more useful than the average “agent benchmark goes up” paper. It treats the reward pipeline as the product, not as a
16 Jun 2026 4 min read
Qwen’s Value Axis Paper Makes Model Confidence Look Less Like a Feeling and More Like a Controllable Feature
ai-models

Qwen’s Value Axis Paper Makes Model Confidence Look Less Like a Feeling and More Like a Controllable Feature

Model confidence is usually treated like weather: visible in the forecast, hard to control, and mostly tolerated until it ruins the deployment. The new Value Axis paper is interesting because it argues confidence is not just an output style or a calibrated probability glued onto the end of generation. In
16 Jun 2026 4 min read
Page 1 of 1
The LGTM © 2026
  • Sign up
Powered by Ghost