Unlearning repeatitive cycles in RL

Dinesh Kumar Banjare, Mithilesh Atulkar, Mitul Kumar Ahirwal, Utilizing machine unlearning for faulty behavior patterns in human decision-making model: A sharded reinforcement learning approach, Engineering Applications of Artificial Intelligence, Volume 181, Part 6, 2026,10.1016/j.engappai.2026.115652.

In this paper, enhance the computational models of human decision-making by proposing a novel sharded reinforcement learning framework based on the principle of machine unlearning. The work proposes a data sharding and selective machine unlearning approach to unlearn faulty behavior patterns. It overcomes the drawbacks of conventional reinforcement learning methods, in which the players tend to play the same deck over and over again, showing a lack of adaptability. These problems are frequently seen in tasks such as the Iowa gambling task. To handle this issue, the concept of machine unlearning is used. The proposed method detects repeated patterns of faulty behavior, such as repeated deck selections. Such patterns are corrected locally at the level of shards, while retaining the last stable expectancy state from the previous learning stage. The mean squared deviation was used to evaluate the model’s performance. The results indicate that the unlearned model learns faster and has better performance in the early phases of learning. The framework supports behavioral adaptability by selectively excluding identified faulty shards while preserving previous valid learning. These findings show the potential of shard-based learning and selective unlearning in sequential decision-making.

Comments are closed.

Post Navigation