Skip to main navigation Skip to search Skip to main content

A reinforcement learning-driven hyper-heuristic algorithm for haul transportation and terminal delivery optimization in two-echelon distribution systems: A case study in GBA

  • Zhi Tang
  • , Ting Qu (Corresponding Author)
  • , Yongheng Zhang
  • , Yanghua Pan
  • , Liqiang Ding
  • , George Q. Huang

Research output: Journal article publicationJournal articleAcademic researchpeer-review

Abstract

The two-echelon distribution system has been increasingly adopted in modern e-commerce logistics. However, as customer service requirements become more diverse, logistics providers must simultaneously accommodate delivery and pickup requests while also satisfying multiple non-overlapping time windows. These additional constraints substantially increase the complexity of balancing service quality and operational costs. To address this challenge, this study extends the classical two-echelon vehicle routing problem and introduces a new variant, called the two-echelon vehicle routing problem with simultaneous pickup and delivery under multiple time windows (2E-VRPSPDMTW). A mixed-integer linear programming (MILP) model is formulated to minimize total operational cost. Given the NP-hard nature of the problem, a Q-learning-based hyper-heuristic algorithm (QLHHA) is developed. The proposed framework first applies a spatiotemporal clustering strategy to allocate customers to satellites, thereby reducing the search space. It then constructs a pool of eight low-level heuristic operators, while a Q-learning mechanism serves as the high-level controller to adaptively select the most appropriate operator. A case study based on real operational data from a cross-border e-commerce logistics company in the Guangdong-Hong Kong-Macau Greater Bay Area (GBA) is conducted to evaluate the method. Extensive test cases and ablation experiment results demonstrate that QLHHA surpasses several state-of-the-art algorithms in both solution quality and stability, achieving up to a 10% reduction in total operational cost. Sensitivity analyses further reveal that, for large-scale demand scenarios, moderately widening the time-window width can substantially reduce operational cost.

Original languageEnglish
Article number131206
Number of pages18
JournalExpert Systems with Applications
Volume309
DOIs
Publication statusPublished - 5 May 2026

Keywords

  • Cross-border logistics
  • Multiple time windows
  • Q-learning based hyper-heuristic
  • Simultaneous pickup and delivery
  • Two-echelon vehicle routing problem

ASJC Scopus subject areas

  • General Engineering
  • Computer Science Applications
  • Artificial Intelligence

Fingerprint

Dive into the research topics of 'A reinforcement learning-driven hyper-heuristic algorithm for haul transportation and terminal delivery optimization in two-echelon distribution systems: A case study in GBA'. Together they form a unique fingerprint.

Cite this