Capability-Specific Degradation Patterns in Quantized Small Language Models

  • Emil Rahimov
Keywords: Code Generation, Edge Artificial Intelligence, Model Quantization, Multilingual Reasoning, Small Language Models

Abstract

Post-training quantization is the default method for deploying small language models (one to four billion parameters) on consumer and edge hardware, yet its effect is usually summarized by a single aggregate accuracy score that can conceal severe failures in individual capabilities. This paper presents a capability-level analysis of four-bit quantization for seven open instruction-tuned small language models drawn from five architecture families: Qwen2.5, Llama-3.2, Gemma-2, Phi-3.5, and SmolLM2. Each model is evaluated at sixteen-bit floating point and at four-bit precision across six capabilities, namely factual knowledge, commonsense reasoning, mathematical reasoning, multilingual mathematical reasoning, code generation, and instruction following, yielding eighty-four controlled evaluations. Instead of reporting only mean accuracy, we construct a per-capability degradation map and test, using Kendall's rank correlation, whether capabilities degrade in a consistent order across architectures. The results show that degradation is strongly capability-specific: multilingual mathematical reasoning and code generation are the most fragile capabilities, with relative losses of up to fifty-seven percent, whereas commonsense reasoning is almost entirely preserved. However, the ordering of degradation is only weakly consistent across architectures, with a mean Kendall's tau of 0.29, indicating that the most and least fragile capabilities are shared but the overall ranking is architecture-dependent. Smaller models degrade more in magnitude. We conclude that small-model quantization should be evaluated per capability rather than in aggregate.

Downloads

Download data is not yet available.

Author Biography

Emil Rahimov

Department of AI and Data Science, Digital Research Lab. Nakhchivan, Azerbaijan.

This is an open access article, licensed under CC-BY-SA

Creative Commons License
Published
        Views : 23
2026-06-24
    Downloads : 13
How to Cite
[1]
E. Rahimov, “Capability-Specific Degradation Patterns in Quantized Small Language Models”, International Journal of Artificial Intelligence, vol. 13, no. 1, pp. 16-23, Jun. 2026.
Section
Articles

References

A. Vaswani et al., “Attention is all you need,” in Advances in Neural Information Processing Systems, 2017, vol. 30, pp. 5998–6008. doi: 10.5555/3295222.3295349.

T. B. Brown et al., “Language models are few-shot learners,” in Advances in Neural Information Processing Systems, 2020, vol. 33, pp. 1877–1901. doi: 10.5555/3495724.3495883.

H. Touvron et al., “LLaMA: Open and efficient foundation language models,” arXiv preprint arXiv:2302.13971, 2023. doi: 10.48550/arXiv.2302.13971.

T. Dettmers, M. Lewis, Y. Belkada, and L. Zettlemoyer, “LLM.int8(): 8-bit matrix multiplication for transformers at scale,” arXiv preprint arXiv:2208.07339, 2022. doi: 10.48550/arXiv.2208.07339.

E. Frantar, S. Ashkboos, T. Hoefler, and D. Alistarh, “GPTQ: Accurate post-training quantization for generative pre-trained transformers,” arXiv preprint arXiv:2210.17323, 2022. doi: 10.48550/arXiv.2210.17323.

J. Lin et al., “AWQ: Activation-aware weight quantization for LLM compression and acceleration,” arXiv preprint arXiv:2306.00978, 2023. doi: 10.48550/arXiv.2306.00978

T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, “QLoRA: Efficient finetuning of quantized LLMs,” arXiv preprint

R. Jin, J. Du, W. Huang, W. Liu, J. Luan, B. Wang, and D. Xiong, “A comprehensive evaluation of quantization strategies for large language models,” in Findings of the Association for Computational Linguistics: ACL 2024, 2024, pp. 12186–12215. doi: 10.18653/v1/2024.findings-acl.726

S. Li et al., “Evaluating quantized large language models,” arXiv preprint arXiv:2402.18158, 2024. doi: 10.48550/arXiv.2402.18158.

K. Marchisio, S. Dash, H. Chen, D. Aumiller, A. Üstün, S. Hooker, and S. Ruder, “How does quantization affect multilingual LLMs?,” in Findings of the Association for Computational Linguistics: EMNLP 2024, 2024, pp. 15928–15947. doi: 10.18653/v1/2024.findings-emnlp.935.

W. Huang et al., “An empirical study of LLaMA3 quantization: From LLMs to MLLMs,” arXiv preprint arXiv:2404.14047, 2024. doi: 10.48550/arXiv.2404.14047.

G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han, “SmoothQuant: Accurate and efficient post-training quantization for large language models,” arXiv preprint arXiv:2211.10438, 2022. doi: 10.48550/arXiv.2211.10438.

Z. Yao, R. Y. Aminabadi, M. Zhang, X. Wu, C. Li, and Y. He, “ZeroQuant: Efficient and affordable post-training quantization for large-scale transformers,” arXiv preprint arXiv:2206.01861, 2022. doi: 10.48550/arXiv.2206.01861.

T. Dettmers and L. Zettlemoyer, “The case for 4-bit precision: k-bit inference scaling laws,” arXiv preprint arXiv:2212.09720, 2022. doi: 10.48550/arXiv.2212.09720.

S. Ashkboos et al., “QuaRot: Outlier-free 4-bit inference in rotated LLMs,” arXiv preprint arXiv:2404.00456, 2024. doi: 10.48550/arXiv.2404.00456.

S. Kim et al., “SqueezeLLM: Dense-and-sparse quantization,” arXiv preprint arXiv:2306.07629, 2023. doi: 10.48550/arXiv.2306.07629.

E. Frantar and D. Alistarh, “SparseGPT: Massive language models can be accurately pruned in a single forward pass,” arXiv preprint arXiv:2301.00774, 2023. doi: 10.48550/arXiv.2301.00774.

B. Jacob et al., “Quantization and training of neural networks for efficient integer-only information,” arXiv preprint arXiv:1712.05877, 2017. doi: 10.48550/arXiv.1712.05877.

S. Han, H. Mao, and W. J. Dally, “Deep Compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding,” arXiv preprint arXiv:1510.00149, 2015. doi: 10.48550/arXiv.1510.00149.

I. Hubara, M. Courbariaux, R. El-Yaniv, and Y. Bengio, “Training neural networks with low precision weights and activations,” arXiv preprint arXiv:1609.07061, 2016. doi: 10.48550/arXiv.1609.07061.

R. Banner, Y. Nahshan, and D. Soudry, “Post-training 4-bit quantization of convolutional networks for rapid-deployment,” arXiv preprint arXiv:1810.05723, 2018. doi: 10.48550/arXiv.1810.05723.

M. Nagel, R. A. Amjad, M. v. Baalen, C. Louizos, and T. Blankevoort, “Up or down? Adaptive rounding for post-training quantization,” in International Conference on Machine Learning (ICML), 2020, pp. 7197–7206.

S. Shen, Z. Yao, A. Gholami, M. Mahoney, and K. Keutzer, “Q-BERT: Hessian based ultra low precision quantization of BERT,” arXiv preprint arXiv:1909.05840, 2019. doi: 10.48550/arXiv.1909.05840.

Y. Bondarenko, M. Nagel, and T. Blankevoort, “Understanding and overcoming the challenges of efficient transformer quantization,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2021, pp. 7947–7958. doi: 10.18653/v1/2021.emnlp-main.627.

X. Zhu, J. Li, Y. Liu, C. Ma, and W. Wang, “A survey on model compression for large language models,” Transactions of the Association for Computational Linguistics, vol. 12, pp. 1556–1577, 2024. doi: 10.1162/tacl_a_00704.

E. J. Hu et al., “LoRA: Low-rank adaptation of large language models,” arXiv preprint arXiv:2106.09685, 2021. doi: 10.48550/arXiv.2106.09685.

A. Yang et al., “Qwen2.5 technical report,” arXiv preprint arXiv:2412.15115, 2024. doi: 10.48550/arXiv.2412.15115.

A. Grattafiori et al., “The Llama 3 herd of models,” arXiv preprint arXiv:2407.21783, 2024. doi: 10.48550/arXiv.2407.21783.

Gemma Team, M. Riviere, et al., “Gemma 2: Improving open language models at a practical size,” arXiv preprint arXiv:2408.00118, 2024. doi: 10.48550/arXiv.2408.00118.

M. Abdin et al., “Phi-3 technical report: A highly capable language model locally on your phone,” arXiv preprint arXiv:2404.14219, 2024. doi: 10.48550/arXiv.2404.14219..

L. B. Allal et al., “SmolLM2: When smol goes big -- Data-centric training of a 1.7B language model,” arXiv preprint arXiv:2502.04353, 2025. doi: 10.48550/arXiv.2502.04353.

L. Gao et al., “A framework for few-shot language model evaluation,” in Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics (EACL), 2023, pp. 270–283.

P. Clark et al., “Think you have solved question answering? Try ARC, the AI2 reasoning challenge,” arXiv preprint arXiv:1803.05457, 2018. doi: 10.48550/arXiv.1803.05457.

R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi, “HellaSwag: Can a machine really finish your sentence?,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019, pp. 4791–4800. doi: 10.18653/v1/P19-1472

K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman, “Training verifiers to solve math word problems,” arXiv preprint arXiv:2110.14168, 2021. doi: 10.48550/arXiv.2110.14168.

F. Shi et al., “Language models are multilingual chain-of-thought reasoners,” arXiv preprint arXiv:2210.03057, 2022. doi: 10.48550/arXiv.2210.03057.

M. Chen et al., “Evaluating large language models trained on code,” arXiv preprint arXiv:2107.03374, 2021. doi: 10.48550/arXiv.2107.03374.

J. Zhou et al., “Instruction-following evaluation for large language models,” arXiv preprint arXiv:2311.07911, 2023. doi: 10.48550/arXiv.2311.07911..

D. Hendrycks et al., “Measuring massive multitask language understanding,” arXiv preprint arXiv:2009.03300, 2020. doi: 10.48550/arXiv.2009.03300.

J. Wei et al., “Chain-of-thought prompting elicits reasoning in large language models,” arXiv preprint arXiv:2201.11903, 2022. doi: 10.48550/arXiv.2201.11903.

W. Kwon et al., “Efficient memory management for large language model serving with pagedattention,” arXiv preprint arXiv:2309.06180, 2023. doi: 10.48550/arXiv.2309.06180.

M. G. Kendall, “A new measure of rank correlation,” Biometrika, vol. 30, no. 1-2, pp. 81–93, 1938. doi: 10.1093/biomet/30.1-2.81.