ORCID
https://orcid.org/0009-0008-7604-249X
Date of Award
Summer 2026
Language
English
Embargo Period
7-30-2026
Document Type
Master's Thesis
Degree Name
Master of Science (MS)
College/School/Department
Department of Information Science and Technology
Program
Information Science
First Advisor
Kimberly A. Cornell
Committee Members
Carol Anne Germain, Brian H. Nussbaum
Keywords
Artificial Intelligence, Large Language Models, Jailbreaking, Cybersecurity, Cybercrime
Subject Categories
Cybersecurity
Abstract
Generative AI (GenAI) and Large Language Models (LLMs) have made large strides in coding task capabilities, with many software developers integrating agentic engineering into their workflow. While GenAI has largely benefited professional software engineers who can automate their work, it has also created room for those with little coding expertise to also create fully fledged programs and applications. It is commonly noted that GenAI is trained with two major goals in mind: to be as helpful as possible, and be as harmless as possible. There exist moments where helpfulness may be prioritized over harmlessness when these goals conflict. LLMs may be jailbroken into fulfilling malicious prompts that safeguards otherwise would have prevented, stemming from the exploitation of helpfulness. This has subsequently allowed threat actors and cybercriminals to utilize GenAI for malicious code generation with the intent to automate cyberattacks. Given that different jailbreak techniques have varying levels of success, this study analyzes the scope of success in malicious code generation that amateur threat actors may produce when jailbreaking various open-source LLMs. This study further analyzes the functionality of the generated malicious scripts to determine whether they function as intended, in order to explore the accuracy of AI-generated malicious code. This allows for exploration in how much human intervention is needed to take AI-generated malicious code into functional malware.
License
This work is licensed under the University at Albany Standard Author Agreement.
Recommended Citation
Capodieci, Noelle, "When Helpfulness Becomes Harmful: Jailbreaking LLMs for Malicious Code Generation" (2026). Electronic Theses & Dissertations (2024 - present). 522.
https://scholarsarchive.library.albany.edu/etd/522