Home
Research
Publications
Education
Research Experience
Projects
Teaching
Skills
Contact
Publications
Interpretable Adaptive Sampling for LLM Test-Time Scaling
An interpretable adaptive sampling method that dynamically allocates inference-time compute for efficient LLM test-time scaling.
Mobina Kashaniyan
,
Ali Jannesari
Posted Aug 4, 2026
Preprint
,
Artificial Intelligence
,
Large Language Models
PDF
DOI
PDF
arXiv
Google Scholar
Research Summary
An Auto-Scaling Approach for Serverless Environments Based on a Multi-Expert Consensus Mechanism
A dependency-aware auto-scaling framework for serverless computing using multi-expert consensus, workload forecasting, ensemble learning, and cost-aware resource allocation.
Mobina Kashaniyan
,
Mehrdad Ashtiani
,
Amirhossein Ghassemi
Posted Jun 23, 2026
Journal Article
PDF
DOI
PDF
Official SAGE Page
Google Scholar
Research Summary
PerfMamba: Performance Analysis and Pruning of Selective State Space Models
PerfMamba analyzes the runtime behavior, resource utilization, memory access patterns, scaling properties, and pruning opportunities of Mamba-1 and Mamba-2 selective state space models.
Abdullah Al Asif
,
Mobina Kashaniyan
,
Sixing Yu
,
Juan Pablo Muñoz
,
Ali Jannesari
Posted Jan 1, 2026
Conference Paper
,
Efficient AI
,
High-Performance Computing
PDF
PDF
Research Summary
LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4
A fully automated LLM-driven AutoML framework for cross-lingual handwritten OCR using closed-loop neural architecture search with GPT-5, GPT-4o, and Claude Sonnet 4 across Arabic, English, and Persian scripts.
Mobina Kashaniyan
,
Amirhossein Ghassemi
,
Nasser Mozayani
Posted Oct 28, 2025
Conference Paper
PDF
DOI
PDF
IEEE Xplore
Google Scholar
Research Summary