Evaluation of hybrid parallelism for scalable training of DenseNet-121 in diabetic retinopathy classification
Abstract
Training large and complex deep learning models is often constrained by GPU memory limitations and prolonged training times. While several parallelism strategies have been proposed, this study specifically evaluates hybrid parallelism—a combination of data parallelism and pipeline parallelism—to address both challenges simultaneously. Using a case study on diabetic retinopathy (DR) classification with the DenseNet-121 architecture, we analyze the trade-off between computational efficiency and memory scalability. Results show that although hybrid parallelism does not yet provide speedup compared to a single-GPU setup—due to communication overhead and pipeline fragmentation—it enables training of large models that exceed the memory capacity of a single GPU. The trained model achieved a validation accuracy of 0.737, a quadratic weighted kappa (QWK) of 0.861, and a weighted F1-score of 0.749. In contrast, pure data parallelism showed a potential speedup of up to 1.9× in scenarios where the model still fits within a single GPU. These findings highlight the critical role of hybrid parallelism in overcoming the memory wall in large-scale model training, though optimization to reduce overhead remains a key challenge.
Keywords
Data parallelism; DenseNet-121; Distributed training; Hybrid parallelism; Pipeline parallelism
Full Text:
PDFDOI: http://doi.org/10.11591/ijeecs.v43.i2.pp662-671
Refbacks
- There are currently no refbacks.

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Indonesian Journal of Electrical Engineering and Computer Science (IJEECS)
p-ISSN: 2502-4752, e-ISSN: 2502-4760
This journal is published by the Institute of Advanced Engineering and Science (IAES).