_nullmirror
BLOG MODEL FINDER
RSS
BLOG MODEL FINDER
  • Launch ETA: August 2026
  • 27 September 2025 4 min read

    LLM Fingerprints v1.3: GLM-4 Judiciary, Summarization Collapse

    The latest nullbench run used the revised judge panel with glm-4-9b as the mid judge. The swap cut family correlation and exposed variance that earlier Qwen-anchored panels smoothed over. The aggregate scores confirm the …

  • 14 September 2025 6 min read

    LLM Fingerprints v1.2: efficient mid tier models, judge and lineup refresh

    Building on the previous nullbench1 methodology and early refinements2, we expanded the pool, kept decoding and scoring fixed, and asked if these additions change routing. The primary harness still scores answers only …

  • 5 September 2025 4 min read

    Retrievability is not discovery

    RAG pipelines retrieve a narrow slice of documents that look closest in vector space to a query, but treating that narrowing as discovery ignores the fact that anything framed outside familiar patterns is never even …

  • 4 September 2025 6 min read

    The Reshaping of the Information Ecosystem

    The way people reach information is undergoing a structural break. For two decades, the web grew on the back of search engines sending users outward: a query produced links, traffic flowed to publishers, and the …

  • 2 September 2025 5 min read

    Agent Systems with SLMs, Workflow Engines and Adaptive Routing

    NVIDIA’s June 2025 paper Small Language Models Are the Future of Agentic AI argues that agent systems run better when built around small language models (SLMs) rather than leaning exclusively on large ones1. A …

  • 31 August 2025 3 min read

    nullbench Fingerprints v1.1: Stability Updates and New Routes

    We revised nullbench with a stability metric that avoids collapse on single spikes and a category set split into more granular writing style and language comprehension tracks. The overall portfolio remains the same1, but …

  • 31 August 2025 5 min read

    Trust Boundaries for LLM Agents

    Large language models capable of acting as agents introduces a new layer of risk. These systems are more than passive text generators but are often equipped with tools and can increasingly get wired into developer …

  • 30 August 2025 8 min read

    nullbench Fingerprints: Starting an Operational Playbook

    Control of execution defines ownership. When a large language model file sits on your machine, it runs the same way until you decide to replace it. No silent updates, no routed experiments, no persona patches. When …

  • 24 August 2025 7 min read

    nullbench: Judge Panel and Methodology

    Evaluation of large language models (LLMs) is often presented as a leaderboard problem, ranking systems by performance on tasks with clear, objective answers. Prior work showed that such leaderboards hide collapse …

  • 23 August 2025 6 min read

    nullbench: Bias Benchmarking for Large Language Models

    Nullbench is a controlled, reproducible benchmarking framework for large language models (LLMs) designed to isolate inherent response tendencies by evaluating models in a zero-context environment. Unlike ubiquitous …

← Newer Page 4 of 5 Older →
_nullmirror

only signal survives

Links

  • Disclaimer
  • Privacy Policy

Connect

  • X/Twitter
  • Email
© 2026 _nullmirror
RSS Source