about

I’m a machine learning scientist in Seattle — Senior ML Scientist and ML tech lead at Microsoft. Eleven plus years of taking models from prototype to production: the architecture around them, the integration into real products, and the evaluation that makes shipping them defensible.

now

At Microsoft I take language models from prototype to production across Microsoft 365 Copilot, and lead evaluation for the Office Product Group Copilots — Word, PowerPoint, Excel, OneNote, Teams. Six strands:

Shipping Copilot features. End-to-end LLM features across Word, Excel, PowerPoint, and other Office applications — from prompt and model choice through to what ships. On Microsoft Teams I fine-tuned and deployed GPT-4o mini models for tool calling, choosing the right tool and arguments from a user’s request, which improved latency and accuracy together and cost less to serve than prompting a larger model for the same decision.

System architecture. I contributed to the Office Domain Specific Language (ODSL) and the Semantic Interpreter that turns a plain-language request into executable Office commands, described in Natural Language Commanding via Program Synthesis. I co-architected the zero-shot ODSL framework — which cut latency and made the system materially easier to maintain and extend — and co-created the LLM orchestration and command-execution layer behind Excel Copilot.

Computer-use agents. Part of the team researching agents that drive Office through its interface rather than its API. The work spans an RL post-training recipe for vision-language GUI grounding that improved click accuracy from a small training set, and a training environment to run full trajectories against live PowerPoint Online — on-demand document cloning plus automated grading, so training can run in parallel at scale. Plus the benchmarks and verification frameworks that come with it, PPT-EVAL and VerificAgent.

Production efficiency. I retrained and deployed the ResNet-based chart-recommendation model behind Recommended Charts and Analyze Data in Excel. We then partnered with Intel to quantize it using Neural Compressor, which produced a large latency improvement and cut serving COGS at Excel’s scale. I led benchmarking across the quantized variants and presented the results to the Office Product Group.

Benchmarks. I took the Office Comprehension Benchmark from an idea to a public release: 304 native Office files, 1,022 questions, 6,719 atomic assertions across 12 industries, with the evaluation methodology, the judging setup, the failure-analysis tooling, and the cross-org wrangling that a release like that actually requires.

Evaluation infrastructure. The platform the Office Copilots run their tests on — test sets, pipelines, and quality dashboards handling thousands of tests a day so product decisions can be made quickly and defensibly.

before

Penske Logistics, data scientist, 2017–2021. Freight rate prediction with XGBoost that cut MAPE from 16% to 4%; a driver-safety model that reduced violations among watchlist drivers by 27%; monthly revenue forecasting across 500+ locations. End to end each time — modeling through deployment and retraining.

Delphi Automotive (Aptiv), senior software engineer, 2012–2015. C and C++ on infotainment audio and speech for GM, Chrysler, Audi, and Great Wall Motors.

education

M.S. Business Analytics and M.B.A., University of Tennessee, Knoxville.

B.Tech. Electronics & Communications Engineering, Vellore Institute of Technology.

contact

firoz@firozshaik.com · linkedin · github