papers
Office Comprehension Benchmark (OCB): A Benchmark for Document Comprehension across Office Documents
preprint · arXiv forthcoming · project lead and first author
304 real Office files, 1,022 questions, 6,719 atomic assertions, 12 industries. Evaluated with multi-model LLM judging and majority voting, calibrated against human baselines. ::: {.paper-links} dataset code
:::
PPT-EVAL: A Benchmark for Computer-Use Agents on PowerPoint Tasks
International Conference on Machine Learning (ICML) 2026
A task benchmark for agents that drive PowerPoint through its interface rather than its API. ::: {.paper-links} code
:::
VerificAgent: Domain-Specific Memory Verification for Scalable Oversight of Aligned Computer-Use Agents
ICML 2025 Workshop on Computer Use Agents
systems I worked on
Not an author on the paper, but I contributed to the Office Domain Specific Language (ODSL) and Semantic Interpreter described in Natural Language Commanding via Program Synthesis — specifically the zero-shot ODSL framework and the orchestration and command-execution layer behind Excel Copilot.