_nullmirror
BLOG MODEL FINDER
RSS
BLOG MODEL FINDER
  • Launch ETA: August 2026
  • 5 May 2026 5 min read

    Re-evaluating Ollama's LLM Performance on Apple Silicon

    Ollama’s local LLM performance on Apple Silicon reveals that only a subset of models delivers practical latency for short-response tasks. The results highlight execution characteristics over theoretical …

  • 5 May 2026 5 min read

    Was there an Actual Shift in AI Model Performance

    In early 2026, developers across the open-source ecosystem observed a sudden improvement in the usefulness of AI-generated bug reports. This article explores whether this shift was due to a leap in model capability or …

  • 12 April 2026 6 min read

    Trendslop: Why LLMs Struggle with Strategy

    A recent study reveals that large language models, when tasked with generating strategic advice, tend to produce generic, consensus-aligned recommendations rather than context-specific insights. This phenomenon, dubbed …

  • 28 February 2026 13 min read

    Embedding Models on Affordable Cloud VMs and Apple Silicon

    A benchmark of CPU-only embedding inference from low cost DigitalOcean droplets to Apple Silicon, focusing on runtime overhead, CPU tier effects, and scaling limits.

  • 8 February 2026 10 min read

    Security Gating as a Control Problem

    A practical approach to securing LLM agents by treating tool invocation as a control boundary. Why fail-closed gates, monotonic risk checks, and adversarial benchmarks are more robust than prompt-only defenses for …

  • 1 February 2026 9 min read

    GPT-OSS-20B Sampling & Prompting for Style Control

    Best practices for sampling parameters and system prompt design to optimize style compliance and throughput with local LLMs like GPT-OSS-20B.

  • 30 January 2026 7 min read

    Execution Is Cheaper, Comprehension Is Not

    LLMs have collapsed the cost of execution, but verification, comprehension, and ownership remain the binding constraints in software engineering and adjacent knowledge work.

  • 24 January 2026 8 min read

    LLM Fingerprints v1.5: Redistribution with 4chan Data

    The latest nullbench run shows no step-function gains in base model capability. Instead, fine-tuning, abliteration, and long-context extensions primarily redistribute behavior rather than raise ceilings. The real …

  • 19 January 2026 9 min read

    Agentic AI Raises The Floor More Than The Ceiling

    AI agents make easy software tasks cheaper but don’t automate hard engineering. This post tries to explain why autonomy claims overreach and where AI actually helps.

  • 21 December 2025 3 min read

    2025 Learnings on LLMs and Software Work

    2025 suggests that the gap between expectations and measured outcomes around LLMs in software engineering is no longer subtle. Large investments and confident narratives imply broad productivity gains or partial task …

← Newer Page 2 of 5 Older →
_nullmirror

only signal survives

Links

  • Disclaimer
  • Privacy Policy

Connect

  • X/Twitter
  • Email
© 2026 _nullmirror
RSS Source