Towards effective offensive security llm agents: Hyperparameter tuning, llm as a judge, and a lightweight ctf benchmark
Minghao Shao,Parameswari Krishnamurthy,Nanda Rani,Kimberly Milner,Haoran Xi,Meet Udesh,Saksham Aggarwal,Venkata Sai Charan Putrevu,Sandeep K Shukla
@inproceedings{bib_Towa_2026, AUTHOR = {Shao, Minghao and Krishnamurthy, Parameswari and Rani, Nanda and Milner, Kimberly and Xi, Haoran and Udesh, Meet and Aggarwal, Saksham and Putrevu, Venkata Sai Charan and Shukla, Sandeep K }, TITLE = {Towards effective offensive security llm agents: Hyperparameter tuning, llm as a judge, and a lightweight ctf benchmark}, BOOKTITLE = {AAAI Conference on Artificial Intelligence}. YEAR = {2026}}
Recent advances in LLM agentic systems have improved the
automation of offensive security tasks, particularly for Capture
the Flag (CTF) challenges. We systematically investigate the
key factors that drive agent success and provide a detailed
recipe for building effective LLM-based offensive security
agents. First, we present CTFJudge, a framework leverag-
ing LLM as a judge to analyze agent trajectories and provide
granular evaluation across CTF solving steps. Second, we pro-
pose a novel metric, CTF Competency Index (CCI) for partial
correctness, revealing how closely agent solutions align with
human-crafted gold standards. Third, we examine how LLM
hyperparameters, namely temperature, top-p, and maximum
token length, influence agent performance and automated cy-
bersecurity task planning. For rapid evaluation, we present
CTFTiny, a curated benchmark of 50 representative CTF
challenges across binary exploitation, web, reverse engineer-
ing, forensics, and cryptography. Our findings identify optimal
multi-agent coordination settings and lay the groundwork for
future LLM agent research in cybersecurity.
Cybersecurity and trust formation in digital payment use behaviour in North India
@inproceedings{bib_Cybe_2026, AUTHOR = {Aljaradat, Aya and Shukla, Sandeep K }, TITLE = {Cybersecurity and trust formation in digital payment use behaviour in North India}, BOOKTITLE = {Discover Sustainability}. YEAR = {2026}}
This study examines key factors associated with digital payment systems use behaviour in India, focusing on effort expectancy, grievance redressal, trust, performance expectancy, perceived cybersecurity risks, and social influence. Additionally, it explores the moderating effects of cybercrime experience, education level, and age on these relationships. Survey data were collected from urban regions in North India, and structural equation modelling was employed to analyse the determinants of use behaviour among active users. The analysis reveals that effort expectancy, grievance redressal, and performance expectancy are positively associated with use behaviour. Trust did not exhibit a significant direct association with use behaviour; however, it is positively associated with performance expectancy and is associated with social influence and perceived cybersecurity risks. Cybercrime experience moderates the association between perceived cybersecurity risks and trust, highlighting heterogeneity in trust-formation pathways across user subgroups. This research offers new perspectives on the relationships among trust, cybersecurity concerns, and digital payment use behaviour. It is underpinned by a unique dataset collected from diverse urban regions in North India, offering a novel perspective on user behaviour and use determinants. The findings underscore the importance of mitigating cybersecurity risks and strengthening trust-related mechanisms to support sustained use, offering practical recommendations for policymakers and digital payment providers.
Automating Organizational Cyber Security Policy
Compliance against Industry Standards using
Agentic AI
Rohit Negi,Soumyo V. Chakraborty,Amit Negi,Sandeep K Shukla
@inproceedings{bib_Auto_2026, AUTHOR = {Negi, Rohit and Chakraborty, Soumyo V. and Negi, Amit and Shukla, Sandeep K }, TITLE = {Automating Organizational Cyber Security Policy
Compliance against Industry Standards using
Agentic AI}, BOOKTITLE = {International Symposium on Digital Forensics and Security}. YEAR = {2026}}
Auditing and compliance management are an in-
tegral part of a cybersecurity management system (CSMS).
However, the frequency of audits and compliance checks is
typically once a year for external audits and twice a year
for internal audits. Under audit and compliance management,
cybersecurity policy documents defined according to normative
references are also reviewed. This review is either performed
at the semantic level, which may miss contextual reasoning, or
manually, which requires significant cognitive effort and has a
very high time complexity. To fill this gap, in this paper we
focus on leveraging Artificial Intelligence (AI) in cybersecurity
management. We propose the Multi-Agent Cybersecurity Policy
Analysis & Validation Workbench (MACAW), an agentic AI
framework that leverages Large Language Models (LLMs) and
Retrieval Augmented Generation (RAG) with planning and
tool-use capabilities to analyze policies, frameworks, standards,
and guidelines. Agents understand contextual dependencies and
generate adaptive responses. In this experiment, LLM with
RAG is used to generate contextual prompts. This results in
an automated multi-agent policy analysis tool for compliance
audit and assessing the compliance of policies with relevant
industry standards for security controls. We benchmark our
approach using two sets of cybersecurity policies, one from
the transportation sector and another from the financial and
banking sector, and security controls from two widely accepted
standards - NIST 800-53 as well as ISO 27002:2013 / ISO
27002:2022. Experimental results demonstrate that our agentic
framework achieves superior performance in contextual accuracy
and automation efficiency, with accuracy (as represented by the
F1 score) in the range of 83% to 99% against a benchmark
of consensus decisions from human cybersecurity audit experts.
This has considerable implications for both productivity and
security applications
CyberExplorer: Benchmarking LLM Offensive Security Capabilities in a Real-World Attacking Simulation Environment
Nanda Rani,Farshad Khorrami,Muhammad Shafique,Ramesh Karri,Kimberly Milner,Minghao Shao,Meet Udeshi,Haoran Xi,Venkata Sai Charan Putrevu,Saksham Aggarwal,Sandeep K Shukla,Prashanth Krishnamurthy
Technical Report, arXiv, 2026
@inproceedings{bib_Cybe_2026, AUTHOR = {Rani, Nanda and Khorrami, Farshad and Shafique, Muhammad and Karri, Ramesh and Milner, Kimberly and Shao, Minghao and Udeshi, Meet and Xi, Haoran and Putrevu, Venkata Sai Charan and Aggarwal, Saksham and Shukla, Sandeep K and Krishnamurthy, Prashanth }, TITLE = {CyberExplorer: Benchmarking LLM Offensive Security Capabilities in a Real-World Attacking Simulation Environment}, BOOKTITLE = {Technical Report}. YEAR = {2026}}
Existing benchmarks for LLM-based offensive security agents use isolated, single-
target setups with a known vulnerable service and fixed objective. They measure
exploitation effectively, but miss how real Capture-the-Flag (CTF) participants
triage unknown surfaces, prioritize targets, and allocate effort under uncertainty.
Current evaluations therefore fail to assess strategic reasoning beyond exploita-
tion alone. To address this, we introduce CTFExplorer, a benchmark suite that
shifts offensive security evaluation toward a multi-target setting, which tests how
agents explore, prioritize, and chain attacks. CTFExplorer deploys 40 web-based
vulnerable services within a single environment, where agents must autonomously
discover, distinguish, and exploit targets without predefined guidance. We also
present a reactive multi-agent setup as a reference agent framework and develop
an agent-agnostic evaluation framework that records structured reasoning traces
for fine-grained assessment. This enables behavioral evaluation beyond binary flag
capture, such as how agents manage target selection, handle failed hypotheses,
coordinate across multiple stages, and extract security intelligence
AuditorMatch: A Data-Driven Decision-Support System for Evaluating and Selecting Cybersecurity Auditors
Rohit Negi,Gargi Sarkar,Sandeep K Shukla
Computers & Security, CoSe, 2026
@inproceedings{bib_Audi_2026, AUTHOR = {Negi, Rohit and Sarkar, Gargi and Shukla, Sandeep K }, TITLE = {AuditorMatch: A Data-Driven Decision-Support System for Evaluating and Selecting Cybersecurity Auditors}, BOOKTITLE = {Computers & Security}. YEAR = {2026}}
As digital infrastructures expand and cyber threats intensify, cybersecurity audits have become a critical assurance mechanism for detecting vulnerabilities, validating control effectiveness, and maintaining national cyber resilience. The quality of these audits is fundamentally dependent on the competence of the auditors / auditing organizations conducting them. In practice, auditee organizations often select an auditor / auditing organization through cost-driven or ad hoc processes. However, selecting the auditor / auditing organization is still a challenge for auditee organization despite the expanded vendor list provided by regulatory body as empaneled auditors. Inclusion of an auditor in the empaneled auditors list establishes eligibility but by itself does not provide any measure of its capability. To address this critical cybersecurity gap, we introduce AuditorMatch, a risk-informed cybersecurity evaluation and selection framework designed to strengthen audit assurance. AuditorMatch transforms heterogeneous CERT-In (The Computer Emergency Response Team established by the Government of India) empanelment records into a structured relational schema and applies a multi-criteria scoring model that incorporates technical capability, scope-specific audit experience, certification maturity, human-resource stability, and tooling diversity, factors that directly impact vulnerability detection quality and control assessment accuracy. Using the 2024 CERT-In empanelment dataset, we demonstrate how AuditorMatch provides a data-driven, defensible basis for selecting auditors whose competencies are aligned with an organization’s cyber risk profile. Beyond the Indian context, the framework provides a generalizable path toward standardizing cybersecurity audit quality, improving transparency in national audit ecosystems, and reinforcing systemic cyber resilience through improved assurance processes.
Modeling Behavioral Signals in Job Scams: A Human-Centered Security Study
Goni Anagha,Vishakha Dasi Agrawal,Gargi Sarkar,Kavita Vemuri,Sandeep K Shukla
Technical Report, arXiv, 2026
@inproceedings{bib_Mode_2026, AUTHOR = {Anagha, Goni and Agrawal, Vishakha Dasi and Sarkar, Gargi and Vemuri, Kavita and Shukla, Sandeep K }, TITLE = {Modeling Behavioral Signals in Job Scams: A Human-Centered Security Study}, BOOKTITLE = {Technical Report}. YEAR = {2026}}
ob scams have emerged as a rapidly growing
form of cybercrime that manipulates human decision-making
processes. Existing countermeasures primarily focus on scam
typologies or post-loss indicators, offering limited support for
early-stage intervention. In this study, we examine how be-
havioral decision signals can be operationalized as computa-
tional features for identifying vulnerability-associated signals
in job fraud. Using anonymous survey data collected from
a university population, we analyze two dominant job scam
pathways: payment-based scams that require upfront fees and
task-based scams that begin with small rewards before escalating
to financial demands. Drawing on behavioral economics, we
operationalize sunk cost influence, urgency/time-pressure cues,
and social proof as measurable behavioral signals, and analyze
their association with payment behavior using exact inference
under sparsity and uncertainty-aware estimation, with social
proof treated as a context-dependent legitimacy cue rather than
a standalone predictor. Our results show that urgency/time-
pressure cues are significantly associated with payment behavior,
consistent with their role as proximal compliance triggers during
escalation. In contrast, opportunity-loss/FOMO cues were not
reliably identifiable under the current operationalization in our
encounter subset, highlighting the importance of measurement
fidelity and cue-definition consistency. We further observe that
emotional tone in victim narratives and selective non-response
to sensitive questions vary systematically with financial loss
and reporting behavior, suggesting that missingness may reflect
a combination of survey fatigue and selective non-disclosure
for sensitive items rather than purely random noise. These
findings suggest that incorporating human behavioral signals into
cybercrime detection and warning systems can improve early-
stage risk assessment without intrusive monitoring
Vulnerabilities in Machine Learning for cybersecurity: Current trends and future research directions
Shantanu Pal,Geeta Yadav,Zahra Jadidi,Ahsan Habib,Md Palash Uddin,Chandan Karmakar,Sandeep K Shukla
Journal of Information Security and Applications, JISA, 2026
@inproceedings{bib_Vuln_2026, AUTHOR = {Pal, Shantanu and Yadav, Geeta and Jadidi, Zahra and Habib, Ahsan and Uddin, Md Palash and Karmakar, Chandan and Shukla, Sandeep K }, TITLE = {Vulnerabilities in Machine Learning for cybersecurity: Current trends and future research directions}, BOOKTITLE = {Journal of Information Security and Applications}. YEAR = {2026}}
AI In Cybersecurity Education--Scalable Agentic CTF Design Principles and Educational Outcomes
Haoran Xi,Sandeep K Shukla,Farshad Khorrami,Alon Hillel Tuch,Muhammad Shafique,Ramesh Karri,Minghao Shao,Kimberly Milner,Venkata Sai Charan Putrevu,Nanda Rani,Meet Udeshi,Parameswari Krishnamurthy,Brendan Dolan Gavitt,Siddharth Garg
Technical Report, arXiv, 2026
@inproceedings{bib_AI_I_2026, AUTHOR = {Xi, Haoran and Shukla, Sandeep K and Khorrami, Farshad and Tuch, Alon Hillel and Shafique, Muhammad and Karri, Ramesh and Shao, Minghao and Milner, Kimberly and Putrevu, Venkata Sai Charan and Rani, Nanda and Udeshi, Meet and Krishnamurthy, Parameswari and Gavitt, Brendan Dolan and Garg, Siddharth }, TITLE = {AI In Cybersecurity Education--Scalable Agentic CTF Design Principles and Educational Outcomes}, BOOKTITLE = {Technical Report}. YEAR = {2026}}
Large language models are rapidly changing how
learners acquire and demonstrate cybersecurity skills. However,
when human–AI collaboration is allowed, educators still lack val-
idated competition designs and evaluation practices that remain
fair and evidence-based. This paper presents a cross-regional
study of LLM-centered Capture-the-Flag competitions built on
the Cyber Security Awareness Week competition system. To un-
derstand how autonomy levels and participants’ knowledge back-
grounds influence problem-solving performance and learning-
related behaviors, we formalize three autonomy levels: human-in-
the-loop, autonomous agent frameworks, and hybrid. To enable
verification, we require traceable submissions including conver-
sation logs, agent trajectories, and agent code. We analyze multi-
region competition data covering an in-class track, a standard
track, and a year-long expert track, each targeting participants
with different knowledge backgrounds. Using data from the 2025
competition, we compare solve performance across autonomy
levels and challenge categories, and observe that autonomous
agent frameworks and hybrid achieve higher completion rates
on challenges requiring iterative testing and tool interactions.
In the in-class track, we classify participants’ agent designs and
find a preference for lightweight, tool-augmented prompting and
reflection-based retries over complex multi-agent architectures.
Our results offer actionable guidance for designing LLM-assisted
cybersecurity competitions as learning technologies, including
autonomy-specific scoring criteria, evidence requirements that
support solution verification, and track structures that improve
accessibility while preserving reliable evaluation and engagement.
SeqShield: A Behavioral Analysis Approach to Uncover Rootkits
Paras Ghodeshwar,Sandeep K Shukla,Anand Handa,Nitesh Kumar
Technical Report, arXiv, 2026
@inproceedings{bib_SeqS_2026, AUTHOR = {Ghodeshwar, Paras and Shukla, Sandeep K and Handa, Anand and Kumar, Nitesh }, TITLE = {SeqShield: A Behavioral Analysis Approach to Uncover Rootkits}, BOOKTITLE = {Technical Report}. YEAR = {2026}}
Rootkits are among the most elusive types of malware, capable of bypassing traditional static analysis methods due to their metamorphic behavior. Signature-based detection techniques struggle against these threats, necessitating a shift toward dynamic analysis approaches. We propose SeqShield, a behavior-based rootkit detection approach designed specifically for the Windows OS, leveraging API call sequences for dynamic behavior analysis. Instead of relying on static signatures, SeqShield examines the execution patterns of API calls, which inherently reflect malicious intent. Analyzing API sequences, we can effectively identify rootkit-like behavior. We also employed a metamorphic code engine to generate 10X mutated variants of rootkits, demonstrating their obfuscation strategies. SeqShield applies n-gram analysis to extract bigram and trigram features from these API call sequences, enabling effective detection of rootkit-like activity. Among the models tested, Random Forest achieves the highest accuracy of 97.27% (bigram) and 96.17% (trigram). To optimize performance and decrease the dimension, we apply feature importance ranking using the Gini Impurity Index, iteratively selecting the most significant features. The optimized lower-dimensional feature matrix significantly enhances detection efficiency without sacrificing accuracy. Using the optimized feature set, our approach achieves 96.72% accuracy for bigrams and 97.81% accuracy for trigrams.
An Automated Framework for Cybersecurity Policy Compliance Assessment Against Security Control Standards
Bikash Saha,Sandeep K Shukla
Technical Report, arXiv, 2026
@inproceedings{bib_An_A_2026, AUTHOR = {Saha, Bikash and Shukla, Sandeep K }, TITLE = {An Automated Framework for Cybersecurity Policy Compliance Assessment Against Security Control Standards}, BOOKTITLE = {Technical Report}. YEAR = {2026}}
Understanding Jamtara cybercriminals: poverty, opportunity, and the rise of organized criminal networks
Gargi Sarkar,Sandeep K Shukla,Venkata Sai Charan Putrevu
Humanities and Social Sciences Communications, HSSC, 2026
@inproceedings{bib_Unde_2026, AUTHOR = {Sarkar, Gargi and Shukla, Sandeep K and Putrevu, Venkata Sai Charan }, TITLE = {Understanding Jamtara cybercriminals: poverty, opportunity, and the rise of organized criminal networks}, BOOKTITLE = {Humanities and Social Sciences Communications}. YEAR = {2026}}
Despite having limited formal education, individuals from India’s Jamtara and New Jamtara regions, as well as several economically disadvantaged districts, have become adept at orchestrating sophisticated cybercrimes by exploiting societal trust and systemic vulnerabilities. Beyond targeting domestic victims, these criminal activities have evolved into transnational operations, affecting individuals worldwide. Cybercrimes have been normalized within their local communities, with criminal enterprises increasingly structured as well-organized, family-run businesses, perpetuating a cycle of intergenerational involvement.
Discovering NFT Rug Pulls: Matching Behavior Patterns
Using Graph Isomorphism Networks
Sandeep K Shukla,Trishie Sharma
ACM Transactions on Internet Technology, ACM-TIT, 2025
@inproceedings{bib_Disc_2025, AUTHOR = {Shukla, Sandeep K and Sharma, Trishie }, TITLE = {Discovering NFT Rug Pulls: Matching Behavior Patterns
Using Graph Isomorphism Networks}, BOOKTITLE = {ACM Transactions on Internet Technology}. YEAR = {2025}}
Amid the surge of Non-Fungible Tokens (NFTs) in blockchain, this study introduces a meticulous methodology focusing on transaction behaviors to unveil rug pulls — a critical issue impacting financial security
and trust in the NFT landscape. Using a Graph Isomorphism Network (GIN) model with 6 behavioral patterns obtained from transaction sequences, we create a “Rug Pull Pattern Matcher” model. We provide a
comprehensive analysis by applying the model on two datasets — creator’s transactions from 50 reputable
NFT projects and 32 reported rug pulls. Our work utilizes automated labeling to categorize addresses and our
analysis reveals several interconnected NFT creator activities. We present an in-depth mapping of fund flows
and creator interactions exposing suspicious behaviors like artificial inflation and intricate network collaborations among creators. The results of our proposed model demonstrate the efficacy of our methodology
with 75.4% accuracy and 85.9% precision on the dataset of reported rug pulls. This work provides comparative analyses of genuine and malicious creator networks to elucidate their structural differences, helping to
identify genuine and potentially fraudulent NFT activities.