_nullmirror
BLOG MODEL FINDER
RSS
BLOG MODEL FINDER
  • Launch ETA: August 2026
  • 18 August 2026 13 min read

    Intent Alignment Under Ambiguity

    Evaluating Muse Glimmer 30B and Qwen3.8 27B on a new benchmark for inferring user intent under ambiguous or incomplete instructions with local inference.

  • 16 August 2026 4 min read

    Our New Find a Cheaper Model Tool

    A practical model-comparison tool for finding lower-cost alternatives with similar measured capability.

  • 11 August 2026 16 min read

    Can LLMs Preserve Judgment When They Don't Know?

    A 143-prompt adversarial benchmark of hallucination and epistemic control. Meta’s Muse Glimmer sets a new high-water mark, but then scores lower when it thinks harder.

  • 22 July 2026 6 min read

    Durable Memory for Agents

    Store factual knowledge in a temporal, provenance-backed claim graph rather than unstructured markdown text files.

  • 20 June 2026 24 min read

    Searching Is Not Knowing

    What eight open models reveal about current-news research, summarization fidelity, and agent design.

  • 15 June 2026 7 min read

    Retrieval Is Not State

    Challenge the assumption that improvements in context size, retrieval quality, model scale, and memory systems will eventually produce something equivalent to long-term memory, understanding, and reliable reasoning. …

  • 14 June 2026 9 min read

    Evaluating Local Models as VSCode Agents

    Evaluating local models on repository investigation and planning tasks and comparing Qwen3.6, Gemma4, GLM-4.7-Flash, and GPT-OSS-20B, across multiple dimensions of repository-agent quality. Our results show that …

  • 14 June 2026 5 min read

    Local Web Research with SearXNG and Crawl4AI

    This article describes a local research stack that combines SearXNG and Crawl4AI to provide current web search and page retrieval for local language models. It explains how the components interact, how they are …

  • 24 May 2026 16 min read

    Local Vision-Language OCR Benchmark

    We benchmark open-weight vision-language models for local document OCR and semantic page reconstruction. The goal is to identify a practical production configuration that balances extraction quality and operational …

  • 9 May 2026 12 min read

    Fast LLM Judging with No-Think Mode

    We explore the use of “no-think” mode for LLM judges in a local evaluation harness. No-think mode suppresses the reasoning phase, which can improve latency and compliance for strict-output tasks. Our …

Page 1 of 5 Older →
_nullmirror

only signal survives

Links

  • Disclaimer
  • Privacy Policy

Connect

  • X/Twitter
  • Email
© 2026 _nullmirror
RSS Source