AI RESEARCH

When Career Data Runs Out: Structured Feature Engineering and Signal Limits for Founder Success Prediction

arXiv CS.LG

ArXi:2604.00339v1 Announce Type: new Predicting startup success from founder career data is hard. The signal is weak, the labels are rare (9%), and most founders who succeed look almost identical to those who fail. We engineer 28 structured features directly from raw JSON fields -- jobs, education, exits -- and combine them with a deterministic rule layer and XGBoost boosted stumps. Our model achieves Val F0.5 = 0.3030, Precision = 0.3333, Recall = 0.2222 -- a +17.7pp improvement over the zero-shot LLM baseline.