The  LGTM
  • Home
  • Agentic Coding
  • Claude Code
  • Codex
Sign in Subscribe
Multimodal Models Are Judging the Outfit More Than You Think
ai-models

Multimodal Models Are Judging the Outfit More Than You Think

A vision-language model that changes its social judgment when the same person puts on different clothes is not merely “seeing context.” It is making a policy decision from style. That is the reason StylisticBias: A Few Human Visual Cues Drive Most Social Biases in MLLMs is worth attention. The paper
20 Jun 2026 4 min read
LLM Judges Can Infect Each Other With Bias — Measure the Network Before You Trust the Score
ai-models

LLM Judges Can Infect Each Other With Bias — Measure the Network Before You Trust the Score

LLM judges were supposed to make evaluation cheaper. Then they became part of the system being evaluated. That is the uncomfortable premise behind Contagion Networks: Evaluator Bias Propagation in Multi-Agent LLM Systems. The paper studies what happens when model agents evaluate each other and update behavior from those evaluations. The
20 Jun 2026 4 min read
Agent Guardrails Need Probabilities, Not Just Boolean Vibes
ai-models

Agent Guardrails Need Probabilities, Not Just Boolean Vibes

The most dangerous fiction in agent security is that a guardrail is a yes/no box. That fiction is convenient. It lets product teams draw neat diagrams: prompt comes in, classifier checks it, agent calls a tool, policy passes or blocks. It also collapses the actual risk model of a
20 Jun 2026 4 min read
A Claude Code Whole-Drive Complaint Exposes the Read-Permission Gap
claude-code

A Claude Code Whole-Drive Complaint Exposes the Read-Permission Gap

The most revealing thing about the latest Claude Code privacy flare-up is not whether one ls command was rude. It is that developers and agent runtimes are still using the same word — “permission” — to mean different things. A new Claude Code issue alleges that a user asked the tool to
20 Jun 2026 5 min read
Claude Code’s Silent Tool-Call Drops Are an Agent Reliability Bug, Not a Parser Glitch
claude-code

Claude Code’s Silent Tool-Call Drops Are an Agent Reliability Bug, Not a Parser Glitch

A coding agent does not have to delete a file to break your workflow. Sometimes it only has to believe it ran a command that never ran. That is the uncomfortable lesson in a fresh Claude Code issue reporting malformed tool-call markup in long Opus-class sessions. The bug sounds, at
20 Jun 2026 5 min read
AWS Trainium Wants to Be a Merchant Chip. NVIDIA’s Moat Is the Software Gravity Around the Chip
nvidia

AWS Trainium Wants to Be a Merchant Chip. NVIDIA’s Moat Is the Software Gravity Around the Chip

AWS selling Trainium racks outside its own cloud would be more than a product-line expansion. It would be Amazon volunteering for a harder exam: proving its AI silicon can survive outside the managed AWS cocoon, where customers expect not just cheaper tokens but debuggable compilers, boring operations, mature frameworks, predictable
19 Jun 2026 5 min read
Agentic RAG Does Not Need Another Framework. It Needs Fewer PCIe Road Trips
nvidia

Agentic RAG Does Not Need Another Framework. It Needs Fewer PCIe Road Trips

Agentic RAG has spent the last year collecting abstractions. Planners, routers, memory layers, tool registries, tracing dashboards, evaluation harnesses — useful pieces, mostly. But Anubhab Banerjee’s CUDA Top-K retrieval experiment points at a less glamorous problem that many teams still under-measure: the retrieval step keeps crossing the PCIe border like
19 Jun 2026 5 min read
Copilot-Authored PRs Now Show Up Under author:@me. Agent Accountability Is Moving Into Search
azure-ai

Copilot-Authored PRs Now Show Up Under author:@me. Agent Accountability Is Moving Into Search

GitHub just made a small search change with a large accountability implication: pull requests opened by Copilot cloud agent on a user’s behalf now appear in author: searches for that user. Search author:@me on github.com/pulls and the results include both PRs you opened yourself and PRs
19 Jun 2026 4 min read
Opus 4.6 Fast Is Leaving Copilot. Treat Model Deprecations Like Runtime Deprecations, Not UI Churn
azure-ai

Opus 4.6 Fast Is Leaving Copilot. Treat Model Deprecations Like Runtime Deprecations, Not UI Churn

Model deprecations used to be chat-product housekeeping. A name disappeared from a dropdown, a slightly newer one took its place, and most users moved on. Coding agents make that framing obsolete. When a model edits code, reviews pull requests, interprets repo instructions, and participates in agent mode, retiring a model
19 Jun 2026 4 min read
Copilot Finally Exposes Per-User AI Credit Burn. Now Teams Can Stop Guessing Which Agent Workflows Are Expensive
azure-ai

Copilot Finally Exposes Per-User AI Credit Burn. Now Teams Can Stop Guessing Which Agent Workflows Are Expensive

Copilot’s credit meter has finally moved from accounting fog to something engineering teams can actually inspect. GitHub added ai_credits_used to the Copilot usage metrics API, giving enterprise administrators and organization owners a per-user view of AI credit consumption across one-day and 28-day user reports. That sounds like
19 Jun 2026 5 min read
xAI’s Remote MCP and Code Execution Docs Show Grok Becoming an Agent Runtime, Not Just a Model Endpoint
xai

xAI’s Remote MCP and Code Execution Docs Show Grok Becoming an Agent Runtime, Not Just a Model Endpoint

xAI’s Remote MCP and Code Execution docs are the kind of platform update that will not trend until somebody’s agent bill, security review, or tool-call log makes it impossible to ignore. That does not make it minor. This is Grok becoming an agent runtime, not merely a model
19 Jun 2026 4 min read
Grok Connectors Turn xAI’s Chatbot Into an Enterprise Data Plane — Which Means the Permissions Now Matter More Than the Demo
xai

Grok Connectors Turn xAI’s Chatbot Into an Enterprise Data Plane — Which Means the Permissions Now Matter More Than the Demo

Grok connectors are not a chatbot feature. They are xAI stepping into the same enterprise swamp every serious AI assistant eventually enters: email, calendars, document stores, CRM, internal tools, and the permissions model nobody wants to debug after the demo goes well. The refreshed xAI docs describe built-in connectors for
19 Jun 2026 4 min read
← Newer Posts Page 13 of 136 Older Posts →
The LGTM © 2026
  • Sign up
Powered by Ghost