I’m building Tenth Man Labs. We put probabilities on geopolitical events for investment, insurance, and risk teams, with the sources behind each number and a dated record of how it moved.
My background is GPU systems. Lately most of my work is on evaluation and reinforcement learning.
Now
-
Tenth Man Labs Founder
–
Most of the technical work is evaluation: scoring forecasts against questions that later resolve, and catching backtests that can see the future. In private beta, covering China and East Asia.
The name comes from an intelligence practice: when nine analysts agree, the tenth has to assume they’re wrong and go looking for why.
Evals and RL
-
eval-power
How many benchmark items it takes to tell two language models apart. Across 395 models, almost no neighboring leaderboard gaps were significant, and small-pilot plans detected differences only 49–79% of the time, not 80%. In a live six-model test, plans committed before collecting answers detected 23 of 27.
-
seed-power
I tested RL study plans on 11,440 fresh training seeds. Plans targeting 80% power detected a difference in 67 of 119 executable comparisons (56.3%). A small pilot did not deliver the detection rate it promised.
-
control-clock
Seconds from a cold Python process to a policy that passes CartPole or Acrobot on a laptop CPU, with every evaluation on the clock. Search over linear policies passed CartPole in a median 1.1 seconds, nearly four times faster than the fastest PPO; on Acrobot, Stable-Baselines3’s tuned PPO was fastest at 10 seconds.
-
Verge Lab
A rule that chooses preference pairs only when quality aspects agree. On HelpSteer2 human preferences, it agreed with 281 of 286 strict preferences (98.25%), against 277 of 282 (98.23%) for an overall-score-gap baseline at matched yield. I found no clear advantage.
-
AlignmentTax
Does instruction tuning trade calibration for truthfulness? Across seven base/instruct pairs on 790 binary TruthfulQA questions, calibration got worse in six, but only four gained accuracy in exchange.
-
BranchPilot
I learned a stopping rule for self-consistency sampling: decide after each answer whether another sample is worth its penalty. It failed the fixed success rule on GSM8K and on an internal 200-problem MATH-500 holdout; simple rules did better.
-
Faultline
A factory simulator where hidden faults look the same until a cheap test tells them apart. Across 375 matched training seeds, ambiguous-only PPO beat random sampling by 10 points but lost to a difficulty curriculum by 6; the hoped-for advantage over both is ruled out.
-
Quantile cycles
I constructed an exact quantile-control operator that alternates forever between the optimal and a worse action, with no fixed point. Lean checks its smallest rational cycle. A separate 32-seed sampled test did not find the specified Huber failure; the hard-operator result is not a QR-DQN training failure.
Systems
-
cpu-decode
A from-scratch C++ int8 decoder for Qwen2.5-0.5B on a laptop CPU. With 2 threads and short context it reaches 69 tokens per second, 83% of the memory-read ceiling and faster than llama.cpp; at 4,096 tokens of context llama.cpp wins, 44 to 32, because my attention is still scalar.
-
attention-numerics
I ran real FlashAttention-3 FP8 attention in every layer on an H100. Rotation raised Qwen2.5-1.5B’s batch exp-CE ratio to 12.56× versus native BF16; centering keys first brought it to 1.004×. The locked predictor transferred to FA3, but produced 118 false alarms for five harmed SageAttention heads.
-
Aperture
Time to first token for a 27B model on one RTX 6000 Ada: under a second at p99 on cold 3,072-token prompts (803 ms), and 1.09 s at 4,096 tokens, which is where the claim stops.
-
HeliosTune
A Thompson-sampling autotuner for Triton matmul kernels, built to test whether timings from cheaper GPUs can warm-start H100 tuning. Transfer didn’t improve this study, and a later audit found that torch.matmul beat all 36 configurations on all 96 stored H100 workloads (1.61× geometric-mean reference/torch ratio). The action set, not the search, is the limit here; L4 and A10 were mixed.
-
Open source
Landed pull requests in PyTorch, Ray, Sentence Transformers, SGLang, and Celery.
In the browser
-
KAN visualizer
I compare Gaussian-edge networks with parameter-matched MLPs and show the exported models’ real activations. MLPs won both two-dimensional tasks in every paired seed; the edge network won the one-dimensional task. These are small synthetic regression tasks.
-
Attention numerics explorable
Explore when rotating queries and keys before low-precision rounding changes attention error. The illustrations and measured results are labelled separately.
-
Eval power calculator
Plan a paired model comparison from the smallest accuracy difference that matters, rather than trusting the gap in a small pilot.
-
Durability debt
Step through a synthetic producer–consumer handoff where a write becomes visible before it survives a crash. Compare publishing first with flushing first.
-
South Side census map
Explore Census block groups clustered by demographic similarity around Chicago’s South Side. These are exploratory partitions, not validated neighborhood boundaries.
-
PDBView
I inspect protein structures in the browser with 3Dmol.js. Load an RCSB PDB ID or drop a local file, click an observed residue to highlight it, measure atom distances, and share a view of a public structure.
-
Terrasim
I built a browser terrarium in Three.js and TypeScript. Pour soil and water, plant a garden, and return to a browser-local save. Its water and plant-care rules are a playable illustration, not a validated ecosystem model.
Research
-
Privacy-Preserving ML Lab Illinois Tech
–
GPU and memory-system acceleration for secure inference with function secret sharing, where the keys for a transformer-scale model can run past 100 GB.
-
Computational Chemistry Lab Illinois Tech
–
Molecular dynamics of the 5-HT2B serotonin receptor for Alzheimer’s drug discovery, and models that predict which signaling pathway a ligand will favor.
-
Department of Biology Illinois Tech
–24
A Python rewrite of the lab’s C tools for finding water molecules in crystal structures, 90% faster, and a protein viewer used by more than 20 researchers.
Background
-
M.S., Artificial Intelligence Illinois Tech
-
B.S., Computer Science and Statistics Illinois Tech
-
Data science intern John Snow Labs
-
Taught programming to more than 50 students in six countries
–