Alp Cetin

San Francisco

I’m building Tenth Man Labs. We put probabilities on geopolitical events for investment, insurance, and risk teams, with the sources behind each number and a dated record of how it moved.

My background is GPU systems. Lately most of my work is on evaluation and reinforcement learning.

Now

  • Tenth Man Labs Founder

    –

    Most of the technical work is evaluation: scoring forecasts against questions that later resolve, and catching backtests that can see the future. In private beta, covering China and East Asia.

    The name comes from an intelligence practice: when nine analysts agree, the tenth has to assume they’re wrong and go looking for why.

Evals and RL

  • eval-power

    How many benchmark items it takes to tell two language models apart. Across 395 models, almost no neighboring leaderboard gaps were significant, and small-pilot plans detected differences only 49–79% of the time, not 80%. In a live six-model test, plans committed before collecting answers detected 23 of 27.

  • seed-power

    I tested RL study plans on 11,440 fresh training seeds. Plans targeting 80% power detected a difference in 67 of 119 executable comparisons (56.3%). A small pilot did not deliver the detection rate it promised.

  • control-clock

    Seconds from a cold Python process to a policy that passes CartPole or Acrobot on a laptop CPU, with every evaluation on the clock. Search over linear policies passed CartPole in a median 1.1 seconds, nearly four times faster than the fastest PPO; on Acrobot, Stable-Baselines3’s tuned PPO was fastest at 10 seconds.

  • Verge Lab

    A rule that chooses preference pairs only when quality aspects agree. On HelpSteer2 human preferences, it agreed with 281 of 286 strict preferences (98.25%), against 277 of 282 (98.23%) for an overall-score-gap baseline at matched yield. I found no clear advantage.

  • AlignmentTax

    Does instruction tuning trade calibration for truthfulness? Across seven base/instruct pairs on 790 binary TruthfulQA questions, calibration got worse in six, but only four gained accuracy in exchange.

  • BranchPilot

    I learned a stopping rule for self-consistency sampling: decide after each answer whether another sample is worth its penalty. It failed the fixed success rule on GSM8K and on an internal 200-problem MATH-500 holdout; simple rules did better.

  • Faultline

    A factory simulator where hidden faults look the same until a cheap test tells them apart. Across 375 matched training seeds, ambiguous-only PPO beat random sampling by 10 points but lost to a difficulty curriculum by 6; the hoped-for advantage over both is ruled out.

  • Quantile cycles

    I constructed an exact quantile-control operator that alternates forever between the optimal and a worse action, with no fixed point. Lean checks its smallest rational cycle. A separate 32-seed sampled test did not find the specified Huber failure; the hard-operator result is not a QR-DQN training failure.

Systems

  • cpu-decode

    A from-scratch C++ int8 decoder for Qwen2.5-0.5B on a laptop CPU. With 2 threads and short context it reaches 69 tokens per second, 83% of the memory-read ceiling and faster than llama.cpp; at 4,096 tokens of context llama.cpp wins, 44 to 32, because my attention is still scalar.

  • attention-numerics

    I ran real FlashAttention-3 FP8 attention in every layer on an H100. Rotation raised Qwen2.5-1.5B’s batch exp-CE ratio to 12.56× versus native BF16; centering keys first brought it to 1.004×. The locked predictor transferred to FA3, but produced 118 false alarms for five harmed SageAttention heads.

  • Aperture

    Time to first token for a 27B model on one RTX 6000 Ada: under a second at p99 on cold 3,072-token prompts (803 ms), and 1.09 s at 4,096 tokens, which is where the claim stops.

  • HeliosTune

    A Thompson-sampling autotuner for Triton matmul kernels, built to test whether timings from cheaper GPUs can warm-start H100 tuning. Transfer didn’t improve this study, and a later audit found that torch.matmul beat all 36 configurations on all 96 stored H100 workloads (1.61× geometric-mean reference/torch ratio). The action set, not the search, is the limit here; L4 and A10 were mixed.

  • Open source

    Landed pull requests in PyTorch, Ray, Sentence Transformers, SGLang, and Celery.

In the browser

  • KAN visualizer

    I compare Gaussian-edge networks with parameter-matched MLPs and show the exported models’ real activations. MLPs won both two-dimensional tasks in every paired seed; the edge network won the one-dimensional task. These are small synthetic regression tasks.

  • Attention numerics explorable

    Explore when rotating queries and keys before low-precision rounding changes attention error. The illustrations and measured results are labelled separately.

  • Eval power calculator

    Plan a paired model comparison from the smallest accuracy difference that matters, rather than trusting the gap in a small pilot.

  • Durability debt

    Step through a synthetic producer–consumer handoff where a write becomes visible before it survives a crash. Compare publishing first with flushing first.

  • South Side census map

    Explore Census block groups clustered by demographic similarity around Chicago’s South Side. These are exploratory partitions, not validated neighborhood boundaries.

  • PDBView

    I inspect protein structures in the browser with 3Dmol.js. Load an RCSB PDB ID or drop a local file, click an observed residue to highlight it, measure atom distances, and share a view of a public structure.

  • Terrasim

    I built a browser terrarium in Three.js and TypeScript. Pour soil and water, plant a garden, and return to a browser-local save. Its water and plant-care rules are a playable illustration, not a validated ecosystem model.

Research

  • Privacy-Preserving ML Lab Illinois Tech

    –

    GPU and memory-system acceleration for secure inference with function secret sharing, where the keys for a transformer-scale model can run past 100 GB.

  • Computational Chemistry Lab Illinois Tech

    –

    Molecular dynamics of the 5-HT2B serotonin receptor for Alzheimer’s drug discovery, and models that predict which signaling pathway a ligand will favor.

  • Department of Biology Illinois Tech

    –24

    A Python rewrite of the lab’s C tools for finding water molecules in crystal structures, 90% faster, and a protein viewer used by more than 20 researchers.

Background

  • M.S., Artificial Intelligence Illinois Tech

  • B.S., Computer Science and Statistics Illinois Tech

  • Data science intern John Snow Labs

  • Taught programming to more than 50 students in six countries

    –

Contact