Skip to main navigation Skip to search Skip to main content

Enhancing Graph Transformer Training through Adaptive Graph Parallelism

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

Graph Transformers, a variant of Graph Neural Networks (GNNs), excel at capturing long-range dependencies but struggle with scalability due to the quadratic complexity of their attention mechanism. We introduce a new training framework that optimizes parallelization strategies based on the graph structure and system configuration. By using sparse operations like sparse matrix-matrix multiplication (SpMM) and sampled dense-dense matrix multiplication (SDDMM), we enhance sparse graph attention speed by up to 3.8x and cut memory use by 77.6% compared to leading frameworks. Additionally, we implement a lightweight reordering strategy for balanced workloads. Our method efficiently processes large-scale graphs with significant scalability improvements, achieving a 5.8x speedup on the ogbn-proteins dataset and a 3.7x speedup on the ogbn-products dataset in distributed training, surpassing previous parallelization methods.

Original languageEnglish (US)
Title of host publicationProceedings - 2025 IEEE International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages1269-1270
Number of pages2
ISBN (Electronic)9798331526436
DOIs
StatePublished - 2025
Event2025 IEEE International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2025 - Milan, Italy
Duration: Jun 3 2025Jun 7 2025

Publication series

NameProceedings - 2025 IEEE International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2025

Conference

Conference2025 IEEE International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2025
Country/TerritoryItaly
CityMilan
Period6/3/256/7/25

All Science Journal Classification (ASJC) codes

  • Artificial Intelligence
  • Computer Networks and Communications
  • Hardware and Architecture

Fingerprint

Dive into the research topics of 'Enhancing Graph Transformer Training through Adaptive Graph Parallelism'. Together they form a unique fingerprint.

Cite this