Optimizing Collaborative Learning: A Hard-Constraint Reinforcement Learning Approach to Fair Team Composition

Herbák, Marcell, Kovásznai, Gergely, Adil, Ali Adil (2026) Optimizing Collaborative Learning: A Hard-Constraint Reinforcement Learning Approach to Fair Team Composition In: Proceedings of the 13th International Conference on Applied Informatics. Eger, Eszterházy Károly Catholic University Líceum Publisher. pp. 133-145.

[thumbnail of ICAI2026-pp133-145.pdf] pdf
ICAI2026-pp133-145.pdf

Download (683kB) [error in script]
Hivatalos webcím (URL): https://doi.org/10.17048/icai.2026.133

Absztrakt (kivonat)

Collaborative learning requires pedagogically balanced teams to function effectively. However, satisfying multiple strict constraints, such as specific gender ratios and minimized intra-group skill variance, transforms team composition into an NP-hard combinatorial optimization problem. Exact solving approaches, such as Optimization Modulo Theories (OMT), suffer from combinatorial explosion, taking hours to evaluate small cohorts (N = 40). We propose a Deep Reinforcement Learning (DRL) framework to resolve this bottleneck. Adapting the Long and Short-Term Constraints (LSTC) architecture, we enforce non-negotiable rules via dynamic action masking and cubic reward shaping. We utilize Deep Q-learning from Demonstrations (DQfD) via replay buffer pre-loading to initialize the policy with valid baselines. Empirical results show that our DRL agent achieves a significant computational speedup over the OMT-baseline while matching its engagement maximization. Furthermore, our approach eliminates the “failure clusters” prevalent in global optimization, improving worst-case team fairness. Robustness testing proves that the policy remains highly deterministic across randomized initializations.

Mű típusa: Könyvrészlet - Book section
Szerző:
Szerző neve
Email
MTMT azonosító
ORCID azonosító
Közreműködés
Herbák, Marcell
NEM RÉSZLETEZETT
NEM RÉSZLETEZETT
NEM RÉSZLETEZETT
Szerző
Kovásznai, Gergely
NEM RÉSZLETEZETT
NEM RÉSZLETEZETT
NEM RÉSZLETEZETT
Szerző
Adil, Ali Adil
NEM RÉSZLETEZETT
NEM RÉSZLETEZETT
NEM RÉSZLETEZETT
Szerző
Kapcsolódó URL-ek:
Kulcsszavak: Collaborative Learning, Team Formation, Deep Reinforcement Learning, Safe RL, Combinatorial Optimization
Nyelv: angol
DOI azonosító: 10.17048/icai.2026.133
Felhasználó: Tibor Gál
Dátum: 22 Szep 2026 07:08
Utolsó módosítás: 22 Szep 2026 07:08
URI: http://publikacio.uni-eszterhazy.hu/id/eprint/9441
Műveletek (bejelentkezés szükséges)
Tétel nézet Tétel nézet