about

I’m a machine learning scientist in Seattle. Eleven years in, currently Senior ML Scientist and ML tech lead at Microsoft.

now

At Microsoft I build LLM-powered features for Microsoft 365 Copilot and lead evaluation across the Office Product Group Copilots — Word, Excel, PowerPoint, OneNote. Three strands:

Benchmarks. I took the Office Comprehension Benchmark from an idea to a public release: 304 native Office files, 1,022 questions, 6,719 atomic assertions across 12 industries, with the evaluation methodology, the judging setup, the failure-analysis tooling, and the cross-org wrangling that a release like that actually requires.

Computer-use agents. An RL post-training recipe for vision-language GUI grounding, and a training environment that runs full trajectories against live PowerPoint Online with on-demand document cloning and automated grading.

Evaluation infrastructure. The platform the Office Copilots run their tests on — test sets, pipelines, and quality dashboards handling thousands of tests a day so product decisions can be made quickly and defensibly.

before

Penske Logistics, data scientist, 2017–2021. Freight rate prediction with XGBoost that cut MAPE from 16% to 4%; a driver-safety model that reduced violations among watchlist drivers by 27%; monthly revenue forecasting across 500+ locations. End to end each time — modeling through deployment and retraining.

Delphi Automotive (Aptiv), senior software engineer, 2012–2015. C and C++ on infotainment audio and speech for GM, Chrysler, Audi, and Great Wall. Led a five-person team and redesigned Chrysler’s text-to-speech module.

education

M.S. Business Analytics and M.B.A., University of Tennessee, Knoxville. B.Tech. Electronics & Communications Engineering, Vellore Institute of Technology.

contact

firoz@firozshaik.com · linkedin · github · scholar · orcid