Deepika Dubey, Geetam Singh Tomar, Dinesh Sharma
The growing adjustment capability of cyber attackers poses hurdles to conventional security systems, as opponent nonstop change their strategies to overcome existing defenses. Conventional approaches, like rule-based firewalls and signature-based intrusion detection systems, based on static assumptions and often unsuccessful by evolving threats. Reinforcement learning (RL) present a dynamic option; however, most prior work considers a single learning agent interacting with a fixed competitor.This paper offers an adversarial reinforcement learning framework in which attacker and defender behave as an autonomous representative within a network shows as a structure of graph. Their communication is behave as a zero-sum Markov game to express the competitive nature of cyber conflicts. The model join the probabilistic intrusion detection, attacker trade-offs between stealth and speed, and resource-constrained protective actions, including concise monitoring and non permanent edge blocking. Through repeated interactions system Both agents utilize independent tabular Q-learning to learn optimal policies. In this paper the experimental results explain that the model reached a high attacker success rate of 97.7% with a close-zero mean reward (0.005), shows fixed and successful learning. The proposed system behave rapidly with an average of one step per episode, and success rates enhance from 96.7% in early stages to 98.7% in later stages, demonstrating consistent performance over time. The results shows the effectiveness of incorporating probabilistic detection and resource-aware procedure in realistic cyber defense scenarios, while providing a actual and interpretable statement for future adversarial cyber security research.