Herbák, Marcell, Kovásznai, Gergely, Adil, Ali Adil (2026) Optimizing Collaborative Learning: A Hard-Constraint Reinforcement Learning Approach to Fair Team Composition In: Proceedings of the 13th International Conference on Applied Informatics. Eger, Eszterházy Károly Catholic University Líceum Publisher. pp. 133-145.
|
pdf
ICAI2026-pp133-145.pdf Download (683kB) [error in script] |
Absztrakt (kivonat)
Collaborative learning requires pedagogically balanced teams to function effectively. However, satisfying multiple strict constraints, such as specific gender ratios and minimized intra-group skill variance, transforms team composition into an NP-hard combinatorial optimization problem. Exact solving approaches, such as Optimization Modulo Theories (OMT), suffer from combinatorial explosion, taking hours to evaluate small cohorts (N = 40). We propose a Deep Reinforcement Learning (DRL) framework to resolve this bottleneck. Adapting the Long and Short-Term Constraints (LSTC) architecture, we enforce non-negotiable rules via dynamic action masking and cubic reward shaping. We utilize Deep Q-learning from Demonstrations (DQfD) via replay buffer pre-loading to initialize the policy with valid baselines. Empirical results show that our DRL agent achieves a significant computational speedup over the OMT-baseline while matching its engagement maximization. Furthermore, our approach eliminates the “failure clusters” prevalent in global optimization, improving worst-case team fairness. Robustness testing proves that the policy remains highly deterministic across randomized initializations.
| Mű típusa: | Könyvrészlet - Book section |
|---|---|
| Szerző: | Szerző neve Email MTMT azonosító ORCID azonosító Közreműködés Herbák, Marcell NEM RÉSZLETEZETT NEM RÉSZLETEZETT NEM RÉSZLETEZETT Szerző Kovásznai, Gergely NEM RÉSZLETEZETT NEM RÉSZLETEZETT NEM RÉSZLETEZETT Szerző Adil, Ali Adil NEM RÉSZLETEZETT NEM RÉSZLETEZETT NEM RÉSZLETEZETT Szerző |
| Kapcsolódó URL-ek: | |
| Kulcsszavak: | Collaborative Learning, Team Formation, Deep Reinforcement Learning, Safe RL, Combinatorial Optimization |
| Nyelv: | angol |
| DOI azonosító: | 10.17048/icai.2026.133 |
| Felhasználó: | Tibor Gál |
| Dátum: | 22 Szep 2026 07:08 |
| Utolsó módosítás: | 22 Szep 2026 07:08 |
| URI: | http://publikacio.uni-eszterhazy.hu/id/eprint/9441 |
![]() |
Tétel nézet |
