Graph Neural Networks (GNNs) have become the model of choice for processing non‑Euclidean data in applications such as robotics, smart‑grid monitoring and real‑time anomaly detection. Deploying them on resource‑constrained edge devices remains a major challenge because of a structural dichotomy inside every GNN layer: an irregular, memory‑bound neighbourhood aggregation phase followed by a dense, compute‑bound feature transformation phase. Conventional edge CPUs and GPUs cannot bridge this gap efficiently and end up memory‑bottlenecked, with their compute logic idle.
This thesis proposes a heterogeneous hardware‑software co‑design framework for low‑precision GNN inference and training at the edge, splitting the workload across the two domains of an adaptive compute acceleration platform (AMD Versal AI Edge). The SGRACE accelerator framework, adapted to the programmable logic, will orchestrate the irregular dataflows, sparse memory accesses and variable 1‑to‑8‑bit quantisation of GNN layers, including graph attention and graph transformer layers. Dense computation will be streamed to the hardened, high‑frequency AI Engine‑ML vector array through the network‑on‑chip and AXI stream interfaces. Exploiting the native sub‑byte vector capabilities of the AI Engines together with arbitrary‑precision spatial pipelines in the fabric opens the way to dynamic precision scaling across layers and phases, both for inference and for on‑device training.
The anticipated results are twice the energy efficiency (TOPS/W) of current edge hardware and deterministic sub‑millisecond latency, delivering scalable, deployment‑ready architectures for next‑generation intelligent edge systems.
Research objectives
SGRACE‑based aggregation engine in programmable logic with arbitrary 1‑to‑8‑bit precision and support for graph attention and graph transformer layers.Dense transformation kernels on the AI Engine‑ML array with sub‑byte vectorisation and a PL‑to‑AIE streaming dataflow.Dynamic precision scaling and on‑device training supported by a hardware‑aware quantisation methodology.Evaluation on Versal AI Edge devices in TOPS/W, latency and accuracy against edge CPU, GPU and FPGA baselines.
Preferred skills
Digital hardware design in SystemVerilog or VHDL and high‑level synthesis; C/C++ and Python with deep learning frameworks (PyTorch); knowledge of neural network quantisation and graph neural networks; experience with AMD Versal,
AI Engines or FPGA accelerator design is a plus. A high level of English is required; Spanish is optional.
Conditions
Full‑time 3‑year doctoral contract at CEIMM‑UPM in Madrid, funded by European and Spanish research projects, with health insurance, a gross salary of 28,000 euros per year in 12 monthly payments and expenses covered for conferences, summer schools and workshops. Starting date is adaptable, the sooner the better. Doctoral students join a highly international team running more than ten European projects, publish in top venues, co‑supervise junior researchers and may teach as assistants for up to two semesters.
The Technical University of Madrid (UPM) is Spain's largest technical university, offering diverse programs, fostering international collaborations, and excelling in research and innovation across multiple disciplines.
Co-Designing Scalable Low-Precision Graph Neural Networks for Inference and Training at the Edge with Heterogeneous Hardware
Full‑time
Madrid, ES
Graph Neural Networks (GNNs) have become the model of choice for processing non‑Euclidean data in applications such as robotics, smart‑grid monitoring and real‑time anomaly detection. Deploying them on resource‑constrained edge devices remains a major challenge because of a structural dichotomy inside every GNN layer: an irregular, memory‑bound neighbourhood aggregation phase followed by a dense, compute‑bound feature transformation phase. Conventional edge CPUs and GPUs cannot bridge this gap efficiently and end up memory‑bottlenecked, with their compute logic idle.
This thesis proposes a heterogeneous hardware‑software co‑design framework for low‑precision GNN inference and training at the edge, splitting the workload across the two domains of an adaptive compute acceleration platform (AMD Versal AI Edge). The SGRACE accelerator framework, adapted to the programmable logic, will orchestrate the irregular dataflows, sparse memory accesses and variable 1‑to‑8‑bit quantisation of GNN layers, including graph attention and graph transformer layers.
Dense computation will be streamed to the hardened, high‑frequency AI Engine‑ML vector array through the network‑on‑chip and AXI stream interfaces. Exploiting the native sub‑byte vector capabilities of the AI Engines together with arbitrary‑precision spatial pipelines in the fabric opens the way to dynamic precision scaling across layers and phases, both for inference and for on‑device training.
The anticipated results are twice the energy efficiency (TOPS/W) of current edge hardware and deterministic sub‑millisecond latency, delivering scalable, deployment‑ready architectures for next‑generation intelligent edge systems.
Research objectives
SGRACE‑based aggregation engine in programmable logic with arbitrary 1‑to‑8‑bit precision and support for graph attention and graph transformer layers.Dense transformation kernels on the AI Engine‑ML array with sub‑byte vectorisation and a PL‑to‑AIE streaming dataflow.Dynamic precision scaling and on‑device training supported by a hardware‑aware quantisation methodology.Evaluation on Versal AI Edge devices in TOPS/W, latency and accuracy against edge CPU, GPU and FPGA baselines.
Preferred skills
Digital hardware design in SystemVerilog or VHDL and high‑level synthesis; C/C++ and Python with deep learning frameworks (PyTorch); knowledge of neural network quantisation and graph neural networks; experience with AMD Versal, AI Engines or FPGA accelerator design is a plus. A high level of English is required; Spanish is optional.
Conditions
Full‑time 3‑year doctoral contract at CEIMM‑UPM in Madrid, funded by European and Spanish research projects, with health insurance, a gross salary of 28,000 euros per year in 12 monthly payments and expenses covered for conferences, summer schools and workshops. Starting date is flexible, the sooner the better. Doctoral students join a highly international team running more than ten European projects, publish in top venues, co‑supervise junior researchers and may teach as assistants for up to two semesters.
The HiPEAC project has received funding from the European Union's Horizon Europe research and innovation funding programme under grant agreement number . Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union. Neither the European Union nor the granting authority can be held responsible for them.
#J-18808-Ljbffr
📌 Co-Designing Scalable Low-Precision Graph Neural Networks for Inference and Training at the Edge with Heterogeneous Hardware (Madrid)
🏢 HiPEAC
📍 Madrid