Why OpenAI's resume screen weights production ML over research credentials
OpenAI's screen scans for shipping ML at scale + safety alignment + research depth. Reported median IC TC ~$555K and acceptance rate <2% reflect a high-bar filter. Generic ML resumes — Kaggle medal-heavy, framework-user shallow — fail at the human read even when keyword-clean.
Three signals matter most for SWE/RE screening: (1) production ML deployment — specific stack (Triton, vLLM, ONNX), specific optimization (quantization, prompt caching, speculative decoding), specific scale (queries/day, p99 latency, GPU count); (2) research depth signaling — from-scratch transformer implementation, paper reproductions, substantive OSS to ML frameworks (PyTorch, HuggingFace, JAX); (3) safety/alignment engagement (for alignment roles) — alignment forum / lesswrong posts, interpretability OSS, paper critiques. Bullets without production + depth signals fail screening at SWE level.