Home
Research
Publications
Education
Research Experience
Projects
Teaching
Skills
Contact
Generation Schedule
Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling
arXiv:2609.19499 — At fixed candidate count N, how candidates are batched into generation calls can change LLM test-time scaling energy and latency by several times on A100 GPUs.
Mobina Kashaniyan
,
Ali Jannesari
Posted Nov 1, 2026
Conference Paper
,
Workshop Paper
,
Preprint
,
Large Language Models
,
High-Performance Computing
,
Energy Efficiency
PDF
DOI
PDF
arXiv
DOI
Google Scholar
Research Summary
Local PDF
Sample Count Is Not Enough (arXiv:2609.19499): Why Candidate-Generation Strategy Matters for LLM Test-Time Scaling Energy and Performance
arXiv:2609.19499 — Candidate count N alone does not define the systems cost of LLM test-time scaling. At fixed N=8, generation schedules like 1×8 vs 8×1 can change A100 energy by about 4.6–4.9× and P95 latency by about 5.8–6.1×.
Mobina Kashaniyan
Posted Sep 16, 2026
Last updated Sep 17, 2026
Large Language Models
,
High-Performance Computing
,
Energy Efficiency
,
Test-Time Scaling
,
Research Summary