<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Neural Network Acceleration | Mobina Kashaniyan</title><link>https://iammobina.github.io/tag/neural-network-acceleration/</link><atom:link href="https://iammobina.github.io/tag/neural-network-acceleration/index.xml" rel="self" type="application/rss+xml"/><description>Neural Network Acceleration</description><generator>Wowchemy (https://wowchemy.com)</generator><language>en-us</language><lastBuildDate>Thu, 01 Jan 2026 00:00:00 +0000</lastBuildDate><image><url>https://iammobina.github.io/media/logo_hu_49df124fc4898e21.png</url><title>Neural Network Acceleration</title><link>https://iammobina.github.io/tag/neural-network-acceleration/</link></image><item><title>PerfMamba: Performance Analysis and Pruning of Selective State Space Models</title><link>https://iammobina.github.io/publication/perfmamba-performance-analysis-and-pruning-of-selective-state-space-models/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://iammobina.github.io/publication/perfmamba-performance-analysis-and-pruning-of-selective-state-space-models/</guid><description>&lt;p&gt;This paper contributes to research on &lt;strong&gt;PerfMamba, Mamba-1, Mamba-2, selective state space models, performance analysis, model pruning, benchmarking, efficient AI, and high-performance computing&lt;/strong&gt; by providing a systematic empirical study of the runtime behavior and optimization opportunities of Mamba-style architectures.&lt;/p&gt;
&lt;p&gt;Recent sequence modeling research has introduced &lt;strong&gt;selective state space models&lt;/strong&gt; as efficient alternatives to Transformer architectures. However, their real-world performance behavior, memory access patterns, I/O characteristics, resource utilization, and scaling properties require deeper analysis for effective deployment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;PerfMamba&lt;/strong&gt; profiles &lt;strong&gt;Mamba-1 and Mamba-2&lt;/strong&gt; across sequence lengths from &lt;strong&gt;64 to 16,384 tokens&lt;/strong&gt;. The study analyzes computation patterns, memory behavior, I/O characteristics, and scaling trends to identify the components that dominate runtime and resource usage.&lt;/p&gt;
&lt;p&gt;Based on these insights, the paper proposes a pruning technique that removes low-activity states within the SSM component. This supports improved throughput and reduced memory usage while maintaining accuracy under moderate pruning.&lt;/p&gt;
&lt;p&gt;This work is relevant to researchers working on &lt;strong&gt;Mamba, selective state space models, state space models, large language models, sequence modeling, model pruning, model compression, runtime profiling, AI systems, neural network acceleration, and hardware-aware optimization&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Keywords:&lt;/strong&gt; PerfMamba, Mamba, Mamba-1, Mamba-2, selective state space models, state space models, SSM, sequence modeling, Transformer alternatives, performance analysis, benchmarking, runtime profiling, resource utilization, memory access patterns, I/O characteristics, scaling analysis, model pruning, state pruning, model compression, efficient AI, high-performance computing, hardware-aware optimization, large language models, AI systems.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Citation:&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Al Asif, A., Kashaniyan, M., Yu, S., Muñoz, J. P., &amp;amp; Jannesari, A. (2026). PerfMamba: Performance analysis and pruning of selective state space models. In &lt;em&gt;International Symposium on Benchmarking, Measuring and Optimization&lt;/em&gt; (pp. 27–44). Springer, Singapore.&lt;/p&gt;</description></item></channel></rss>