Research Projects

Interactive Demos & Experiments

Autoresearch in New Domains

A Custom CPU Inference Engine in Nearly 300 Experiments

Karpathy's autoresearch loop, pointed at a systems problem instead of a training run. Coding agents built and tuned a C++ inference engine for a 180B mixture-of-experts model on two ten-year-old Dell servers: experts pinned to four memory domains, activations exchanged over RDMA, speculative decoding with the model's own MTP head. Every change was pre-registered, checked bit-exact against a reference, and kept only if it beat the measured noise floor.

Results:

  • 1.45 → 33 tok/s single-stream decode on Qwen3.8-Flash-Next (180B total, 6B active, 4-bit), no GPU
  • Stock llama.cpp on the same model: 5.3 tok/s
  • Nearly 300 experiments, about 1,300 commits, nine days
  • The same loop tuned the fan controller and checked the acoustic enclosure design

Interactive Demos

RookWorld-LM Reasoning & World Model Demo

Experience transparent reasoning with ROOK-LM and RookWorld-LM. Watch the models think step-by-step through chess positions with streaming chain-of-thought visualization.

Key Achievements:

  • 🏆 32.1% Checkmate-in-One - outperforms ChessGPT-Base (26.5%) with 24x fewer parameters
  • ChessGPT: 3B params, NeurIPS'23 dataset award (Feng et al.)
  • 99.9% environment simulation accuracy
  • Self-play without external engines

ROOK-CLF-9M Multi-Component Analysis Demo

Comprehensive evaluation platform for the 9M parameter chess language model featuring interactive analysis, professional benchmarking, and attention visualization. Explore strategic reasoning capabilities through three specialized interfaces.

Features:

  • Selfplay: Interactive chess board with real-time move analysis and autoplay
  • Benchmark: Professional evaluation against research datasets (ChessBench, BIG-bench, Lichess)
  • Interpretability (WIP 🚧): Attention rollout heatmaps and early logit lens visualization
  • Performance: WebGPU acceleration with IndexedDB model caching

Research Benchmarks:

  • 49% action accuracy (ChessBench)
  • 57% checkmate-in-one accuracy (BIG-bench)
  • Real-time evaluation with performance metrics
  • Authentic research methodologies and datasets

Publications

ROOK: Reasoning Over Organized Knowledge

LAION research note detailing the development of language models for strategic reasoning through chess, including architectural innovations and training methodologies.

Datasets & Models

Datasets

Models