A deep reinforcement learning method for container drayage transportation considering customer pairs

Clicks: 14
ID: 313753
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal
Star

Ranked #5 of 30 articles by views in journal of computational design and engineering

Most read Least read

Bar heights use a square-root scale.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
Abstract Container drayage transportation serves as a critical link in global supply chains, yet truck capacity constraints and the complex interplay of multi-customer requirements often compromise drayage efficiency. These factors collectively increase fuel consumption and operational costs, posing significant challenges for logistics optimization. To address these issues, this article investigates a container drayage problem with customer pairs, where each pickup node corresponds to a delivery node. The optimization aims to minimize the trucks’ total fuel consumption. A mixed-integer nonlinear programming model is formulated on a graph-based representation to capture the coupling between task dependencies and truck states. To reduce computational complexity, we linearize the model by introducing several auxiliary variables. Recognizing the exponential growth of solution space in large-scale scenarios, we propose a deep reinforcement learning (DRL) method that integrates a Markov decision process, policy gradient optimization, and an attention mechanism. The method features a sequential decision-making system with an enhanced attention mechanism, a carefully designed cumulative reward function, and tailored training strategies. Specifically, the encoder efficiently extracts task features from depot, pickup, and delivery nodes, while the decoder optimizes feature fusion to guide task selection. Importantly, the model explicitly incorporates symmetry between customer pairs in both the encoder and decoder, thereby improving solution quality. Extensive experiments validate that the mathematical model, solved via Gurobi, obtains optimal solutions for small-scale instances within 1900 seconds, while the proposed DRL method achieves the same optimal solutions within 2700 seconds. For medium- and large-scale instances, DRL outperforms Gurobi, simulated annealing, and large neighborhood search, consistently delivering superior solutions within acceptable computation time, demonstrating strong generalization and robustness. Ablation studies further confirm the individual contributions of the encoder, decoder, and training strategy, with the full model achieving the best performance. These results underscore the potential of DRL as an effective tool for sustainable container drayage optimization.
Reference Key
openalex_W7160908413 Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors Chao Huang, Yinan Cui, Xiaoyang Zhou, Boyang Qu, Li Yan, Hui Zhang
Journal journal of computational design and engineering
Year 2026
DOI
10.1093/jcde/qwag047
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.