Research

Everything
we published.

Specifications, papers, benchmarks and engineering notes. Open licences, real numbers, and the two results that went against us. Those are the ones worth reading first.

Specification

4 September 2026

Open Brand Definition 3.0.2

An open standard for managing brand truth in AI systems. Three layers: Foundation, Governed Context, Automated Production. Foundation is enough for most cases, everything above it is optional.

openbranddefinition.org →
  • 1,079 conformance tests, zero failures
  • CC BY 4.0 for spec and docs, Apache 2.0 for schemas and suite
  • Normative spec, quickstart, 245 KB release package
  • First published 22 July 2026

Paper

Published, 23 pages

Memory utility for retrieval augmented models

A method for measuring whether a retrieved memory actually earned its place: remove it, re-run the task, measure the damage. Counterfactual ablation instead of similarity scores, LLM judges or rank labels, all of which reason in a circle.

supabrain.space →
  • Utility signal correlates with cosine similarity at −0.024 to +0.161
  • The LoRA specialist we trained on those labels lost to plain BM25 on most datasets
  • We published the result and the data anyway
  • Code on GitHub

Benchmark

Open source

trashfire, AI code review benchmark

42 intentionally broken projects across 30+ languages with encrypted ground truth. Six weighted dimensions: security 35, cross-module bugs 25, logic 20, performance 10, best practice 5, code smells 5. Point any file-reading agent at it and get a comparable score.

trashfire.io →
  • 42 projects, 30+ languages, 4,200+ planted bugs
  • Three tiers: Focused, Standard, Ultimate
  • Automated scoring against encrypted ground truth

Engineering note

Build in progress

supabrain, context compilation under a budget

The problem is not longer context. It is deciding what deserves context. Four primitives: persist, index, compile, verify. Every passage carries a file, a line span and a content hash, so an agent can be held to what it read.

supabrain.dev →
  • 61,747 passages: first search 12.4 s in memory, 0.083 s persisted
  • Memory ~650 MB down to 140 MB
  • Local SQLite, read-only tools over stdio JSON-RPC
  • Local and trusted-local only. No hosted runtime, no production claim

Product evaluation

Three week evaluation

supaskills, testing our own catalogue

We ran our own skill catalogue against frontier models for three weeks to find out whether it does what our marketing said. It does not, not in the way we claimed. The catalogue still ships, the claim does not.

supaskills.ai →
  • 1,307 skills, 6 domains, quality scored on 6 dimensions
  • REST API and MCP, live
  • Published finding: "on frontier models, skills don't help the way we claimed"

Next

Want this applied
to your systems?