Jiacheng Shen, Lihan Feng
In human decision-making tasks, individuals learn through trials and prediction errors. When individuals learn the task, some are more influenced by good outcomes, while others weigh bad outcomes more heavily. Such confirmation bias can lead to different learning effects. In this study, we propose a new algorithm in Deep Reinforcement Learning, CM-DQN, which applies the idea of different update strategies for positive or negative prediction errors, to simulate the human decision-making process when the task’s states are continuous while the actions are discrete. We test CM-DQN in the Lunar Lander environment with confirmatory, disconfirmatory bias and non-bias to observe the learning effects. Moreover, we apply the confirmation model in a multi-armed bandit problem, which utilizes the same idea with our proposed algorithm, as a contrast experiment to algorithmically simulate the impact of different confirmation bias in the decision-making process. In both experiments, confirmatory bias indicates a better learning effect.
@article{b7786f3f-b37e-4d73-af7f-7e0d544fbac8,
title={31 (4)},
author={Jiacheng Shen and Lihan Feng},
year={2026},
language={en}
}TY - JOUR TI - 31 (4) AU - Jiacheng Shen AU - Lihan Feng PY - 2026 LA - en ER -
Unknown, Unknown
This chapter discusses metal casting processes, highlighting the diversity and common characteristics among them. The objective is to elucidate the fu
Unknown, Unknown
This study focuses on the fundamental characteristics of solid iron, which is predominantly composed of iron atoms and provides a basis for understand
Ir. Méshac KIME ILUNGA
Ce document traite des procédés métallurgiques spéciaux, en mettant particulièrement l'accent sur l'extraction liquide-liquide, un processus mis au po
Copper solvent extraction units at large hydrometallurgical plants face constraints in metal recovery, phase disengagement, and reagent consumption, d