The increasing adoption of Artificial Intelligence (AI) in edge and on-board systems, particularly in space applications, has intensified the need for efficient and reliable hardware acceleration solutions. In this context, Field-Programmable Gate Arrays (FPGAs) represent a particularly attractive solution due to their reconfigurability, energy efficiency, and balanced performance. However, the landscape of FPGA-based AI accelerators is characterized by a high degree of architectural heterogeneity and by fundamentally different design paradigms, making fair comparison and informed platform selection a challenging task. This thesis addresses the problem of evaluating and comparing FPGA-based AI acceleration platforms for edge and space-oriented inference. Rather than focusing exclusively on peak performance, the work emphasizes a broader set of criteria that are critical in space systems, including architectural efficiency, energy consumption, predictability, flexibility, portability, and suitability for deployment in radiation-prone and resource-constrained environments. To this end, the thesis introduces a structured benchmarking framework that combines quantitative performance and efficiency metrics with qualitative evaluation criteria. The proposed methodology is designed to mitigate platform-dependent biases and to enable a more meaningful comparison across heterogeneous architectures. The framework is applied to a representative set of solutions spanning different design approaches, including programmable overlay accelerators and network-specific architectures generated through automated or semi-automated toolchains. A central contribution of this work is the in-depth analysis of GPU@SAT, a soft Graphics Processing Unit (GPU) accelerator explicitly conceived for space applications. Through direct access to the Register Transfer Level (RTL) implementation, GPU@SAT is studied at an architectural level, and several enhancements are introduced to improve compliance with the OpenCL execution and memory models. In parallel, a complete development environment, referred to as the GPU@SAT DevKit, is designed and implemented to enable systematic experimentation and reproducible deployment on real hardware. This effort transforms GPU@SAT from a simulation-oriented research prototype into an experimentally evaluable platform. In addition to GPU@SAT, the thesis investigates a streaming-oriented acceleration approach based on hls4ml and OmpSs-2@FPGA, demonstrating how highly specialized architectures can achieve excellent performance and energy efficiency for selected workloads. Commercial and vendor-supported solutions, including Vitis AI and VectorBlox Software Development Kit (SDK), are also evaluated, together with the automated RTL generation framework FPG-AI, providing a broad view of the current design space. The experimental evaluation is carried out using two representative Neural Network (NN) workloads, including a space-oriented model, 1D-Justo-LiuNet, which is deployed and measured end-to-end on FPGA-based accelerators for the first time. The results show that no single platform consistently outperforms the others across all metrics. Instead, each architectural approach exhibits distinct strengths and limitations, reflecting inherent trade-offs between flexibility, efficiency, and verifiability. Overall, this thesis does not aim to promote a single acceleration solution. Rather, it provides a structured methodology and an extensive experimental study that clarifies how different FPGA-based AI acceleration strategies align with specific application requirements and system-level constraints, with particular emphasis on space-oriented on-board inference scenarios.
Comparative Analysis of FPGA-based Accelerators for Edge AI in Space Applications
TODARO, GIOVANNI
2026
Abstract
The increasing adoption of Artificial Intelligence (AI) in edge and on-board systems, particularly in space applications, has intensified the need for efficient and reliable hardware acceleration solutions. In this context, Field-Programmable Gate Arrays (FPGAs) represent a particularly attractive solution due to their reconfigurability, energy efficiency, and balanced performance. However, the landscape of FPGA-based AI accelerators is characterized by a high degree of architectural heterogeneity and by fundamentally different design paradigms, making fair comparison and informed platform selection a challenging task. This thesis addresses the problem of evaluating and comparing FPGA-based AI acceleration platforms for edge and space-oriented inference. Rather than focusing exclusively on peak performance, the work emphasizes a broader set of criteria that are critical in space systems, including architectural efficiency, energy consumption, predictability, flexibility, portability, and suitability for deployment in radiation-prone and resource-constrained environments. To this end, the thesis introduces a structured benchmarking framework that combines quantitative performance and efficiency metrics with qualitative evaluation criteria. The proposed methodology is designed to mitigate platform-dependent biases and to enable a more meaningful comparison across heterogeneous architectures. The framework is applied to a representative set of solutions spanning different design approaches, including programmable overlay accelerators and network-specific architectures generated through automated or semi-automated toolchains. A central contribution of this work is the in-depth analysis of GPU@SAT, a soft Graphics Processing Unit (GPU) accelerator explicitly conceived for space applications. Through direct access to the Register Transfer Level (RTL) implementation, GPU@SAT is studied at an architectural level, and several enhancements are introduced to improve compliance with the OpenCL execution and memory models. In parallel, a complete development environment, referred to as the GPU@SAT DevKit, is designed and implemented to enable systematic experimentation and reproducible deployment on real hardware. This effort transforms GPU@SAT from a simulation-oriented research prototype into an experimentally evaluable platform. In addition to GPU@SAT, the thesis investigates a streaming-oriented acceleration approach based on hls4ml and OmpSs-2@FPGA, demonstrating how highly specialized architectures can achieve excellent performance and energy efficiency for selected workloads. Commercial and vendor-supported solutions, including Vitis AI and VectorBlox Software Development Kit (SDK), are also evaluated, together with the automated RTL generation framework FPG-AI, providing a broad view of the current design space. The experimental evaluation is carried out using two representative Neural Network (NN) workloads, including a space-oriented model, 1D-Justo-LiuNet, which is deployed and measured end-to-end on FPGA-based accelerators for the first time. The results show that no single platform consistently outperforms the others across all metrics. Instead, each architectural approach exhibits distinct strengths and limitations, reflecting inherent trade-offs between flexibility, efficiency, and verifiability. Overall, this thesis does not aim to promote a single acceleration solution. Rather, it provides a structured methodology and an extensive experimental study that clarifies how different FPGA-based AI acceleration strategies align with specific application requirements and system-level constraints, with particular emphasis on space-oriented on-board inference scenarios.| File | Dimensione | Formato | |
|---|---|---|---|
|
phd_thesis_TODARO.pdf
embargo fino al 25/05/2029
Licenza:
Tutti i diritti riservati
Dimensione
11.46 MB
Formato
Adobe PDF
|
11.46 MB | Adobe PDF |
I documenti in UNITESI sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/20.500.14242/367834
URN:NBN:IT:UNIPI-367834