Trivedi et al. Res. Trends Int. J. Technol. Innov., July - September 2026, 1 (3) : 1-8
1Department of Computer Science and Engineering, Indian Institute of Technology Kharagpur, Kharagpur, India; 2Department of Electronics Engineering, College of Engineering Pune, Pune, India
Article History
Accepted : 16 Jun 2026
Published : 10 Jul 2026
Publication Issue
Volume 1, Issue 3
July - September 2026
Page Number1–8
Energy consumption in battery-powered embedded systems is strongly influenced by compiler-level optimisation choices that are often selected using desktop-oriented heuristics ill-suited to microcontroller architectures. This paper benchmarks eight GCC and LLVM optimisation flag combinations on an ARM Cortex-M4 platform across five representative embedded workloads, measuring both execution time and energy via a precision current-sense circuit. Loop-unrolling combined with function inlining reduced energy consumption by up to 22 percent relative to -O2 defaults, though with a 6 percent code-size increase that may be prohibitive on flash-constrained devices.
Keywords - compiler optimization, embedded systems, energy efficiency, ARM Cortex-M, LLVM
Compiler optimisation levels such as -O2 and -O3 are tuned primarily for execution speed on desktop-class processors, and their energy implications on resource-constrained microcontrollers, where memory access and clock-gating behaviour differ substantially, remain comparatively under-studied.
Five representative embedded workloads, including a digital filter, a matrix multiply kernel and a JSON parser, were compiled under eight combinations of loop unrolling, function inlining, and vectorisation flags using both GCC 12 and LLVM 16 targeting an ARM Cortex-M4 development board, with energy measured via an INA219 current-sense IC sampling at 1kHz during execution.
Combining aggressive loop unrolling with function inlining reduced measured energy consumption by up to 22 percent relative to the -O2 baseline across the five workloads, with the matrix multiply kernel showing the largest gain, though the same configuration increased compiled code size by an average of 6 percent, a relevant tradeoff for flash-constrained parts.
Energy-aware compiler flag selection can yield meaningful battery-life improvements on microcontroller targets, but code-size tradeoffs must be evaluated per deployment. Future work will explore profile-guided optimisation informed directly by energy measurements.
[1] Tiwari V. et al., Power analysis of embedded software, ISLPED, 1994. [2] Lattner C. and Adve V., LLVM: A compilation framework, CGO, 2004.
© 2026 The Author(s). Published by IJEIA Editorial Office. This is an open access article under the Creative Commons Attribution 4.0 International License (CC BY 4.0).
Devansh Trivedi, Ishaan Mehra (2026). A Comparative Study of Compiler Optimization Techniques for Energy-Efficient Embedded Software. IJEIA, 1(3), 1-8.