AI Security and Trustworthy AI (4280/6280)
Course Logistics
Course Code: 4280/6280
Semester/Year: Spring 2025
Credit Hours: 3
Instructor: Long Cheng
Email: lcheng2@clemson.edu
Office: McAdams 100
Office Hours: TBA
Teaching Team: None (instructor-led)
Meeting Times:
Mondays and Wednesdays, 2:00–3:30 PM
Location: TBA
Modality: In-person or online (TBA)
Course Management System: Canvas
Course Description
This advanced, hands-on course is organized around three pillars at the intersection of artificial intelligence and cybersecurity:
- AI for Security — using AI/ML to improve cybersecurity (network intrusion and anomaly detection, malware detection and classification, phishing/fraud detection, and AI-assisted security operations).
- Security of AI — protecting AI systems from attack (adversarial examples, data poisoning and backdoors, model extraction/inversion and membership inference, and the security of large language models).
- Trustworthy AI — ensuring AI systems are safe, fair, robust, transparent, and accountable (interpretability, fairness, privacy-preserving ML, robustness certification and formal verification, and AI governance).
Students will learn to apply machine learning to defensive security problems, to identify and model threats to AI systems and implement attacks and defenses, and to audit and improve AI systems for fairness, interpretability, privacy, robustness, and accountability. Through programming assignments, case studies, and a sustained capstone project, students develop the technical and analytical skills to build AI systems that are effective for security, secure against attack, and trustworthy in deployment.
Prerequisites
Undergraduate (senior-level) students: CS 301 or permission of instructor.
Graduate students: No prerequisites.
Learning Objectives
Upon successful completion of this course, students will be able to:
- Apply machine learning to cybersecurity problems (AI for Security) — build and evaluate detectors for network intrusion, malware, and phishing/fraud, accounting for class imbalance, concept drift, and operational constraints.
- Design and execute threat models for AI systems, identifying attack vectors, adversaries, capabilities, and assets at risk across the ML lifecycle.
- Implement adversarial attacks and defenses (Security of AI) using Python and deep learning frameworks (TensorFlow/PyTorch), including evasion, poisoning/backdoors, and model-extraction/privacy attacks.
- Apply interpretability and explainability techniques (e.g., LIME, SHAP, attention, Integrated Gradients) to understand and audit model decisions.
- Conduct fairness audits, detect and quantify bias in AI systems, and implement fairness mitigation strategies.
- Design and implement privacy-preserving techniques such as differential privacy (DP-SGD), federated learning, and secure aggregation.
- Verify robustness and reason about AI governance using formal verification/certification techniques and transparency/accountability practices (model cards, audits, regulatory frameworks).
- Evaluate AI holistically across the three pillars, balancing detection effectiveness, security against attack, and trustworthiness trade-offs.
- Architect and present a capstone project that integrates multiple pillars on a real-world or benchmark AI system.
Topics Covered
Organized around the course’s three pillars:
Pillar 1 — AI for Security
- Foundations of Machine Learning for Cybersecurity
- Network Intrusion and Anomaly Detection
- Malware Detection and Classification
- Phishing/Fraud Detection and AI for Security Operations
Pillar 2 — Security of AI
- Threat Modeling for ML Systems and Adversarial Examples
- Advanced Evasion, Transferability, Data Poisoning, and Backdoors
- Model Extraction, Inversion, Membership Inference, and LLM Security
- Adversarial Defenses and Robustness
Pillar 3 — Trustworthy AI
- Interpretability and Explainability
- Fairness: Definitions, Metrics, Detection, and Mitigation
- Privacy-Preserving Machine Learning
- Robustness Certification, Formal Verification, and Safety
- Transparency, Accountability, and AI Governance
Textbook and Readings
No required textbook. The instructor will provide or recommend open-source resources, academic papers, tool documentation, and web-based tutorials as needed for each unit. Links and materials will be shared via Canvas.
Time Expectations
In-class: 3 hours per week (two 90-minute sessions)
Independent work: 6–9 hours per week (programming assignments, readings, case studies, capstone development)
Total: 9–12 hours per week
Grade Breakdown
| Component | Percentage |
|---|---|
| Homework and Labs (8–10 assignments) | 35% |
| Quizzes and Participation | 10% |
| Midterm Project | 20% |
| Individual Capstone Project | 35% |
| Total | 100% |
Assessment Types and Descriptions
Homework and Labs (35%)
Students will complete 8–10 programming assignments and lab exercises over the course of 17 weeks. Each assignment involves writing Python code using TensorFlow or PyTorch to implement concepts from the lecture (e.g., adversarial attack methods, fairness audits, interpretability analysis, privacy mechanisms). Assignments are cumulative and scaffold toward the capstone. Case studies and security audits are embedded within these coding projects—students do not submit separate written audit reports but instead document findings and code in a Jupyter notebook or GitHub repository. Each assignment is graded on correctness of implementation, code quality, documentation, and adherence to submission standards. Late work is penalized 10% per day for up to 48 hours; work submitted more than 48 hours late receives no credit unless prior arrangements are made.
Quizzes and Participation (10%)
Quizzes are short, low-stakes assessments administered in-class or via Canvas to gauge comprehension of key concepts (threat modeling terminology, fairness metrics, privacy definitions, etc.). Participation includes active engagement in discussions, office hours attendance, peer code review, and constructive collaboration on lab work. Together, quizzes and participation form a single 10% component; the instructor will use a combination of quiz scores and qualitative participation observation to assign this grade, aiming to reward consistent engagement in a hands-on course.
Midterm Project (20%)
In weeks 8–9, students undertake a structured midterm project in which they select an AI system (real or synthetic), build a threat model, and execute one or more adversarial attacks or implement a defense mechanism. The project culminates in a written report (5–10 pages) and a short in-class presentation (5–7 minutes). The report should cover the threat model, attack/defense methodology, experimental results, and insights. Grading is based on technical depth, correctness, clarity of presentation, and integration of threat modeling concepts.
Individual Capstone Project (35%)
Beginning in week 12 and concluding in week 17, each student undertakes a sustained, individual capstone project that integrates across the course’s three pillars (e.g., build an ML-based detector and then secure and audit it, or combine adversarial robustness + fairness, or privacy + interpretability + certification). Students select or are assigned a real-world or benchmark AI system and conduct a comprehensive audit, design and implement improvements, and evaluate the system holistically. The capstone culminates in two deliverables:
- Written Report (15–20 pages): A technical document detailing the system under study, threat models, experimental findings, implementations, and trade-offs between trustworthiness dimensions.
- Presentation (10–15 minutes): An oral presentation with slides and/or live demo, delivered during the final exam period or a designated capstone presentation session, followed by Q&A.
Grading criteria include technical rigor, novelty of approach, integration of course concepts, clarity of documentation and presentation, and reproducibility of results. Interim milestones (project proposal in week 12, progress report in week 14) are embedded in the schedule to ensure sustained progress and timely feedback.
Attendance and Participation Expectations
Regular attendance is expected. In-person or synchronous online attendance allows you to engage with lectures, participate in discussions, and collaborate on lab work. If you must miss class due to illness, emergency, or other valid reason, inform the instructor as soon as possible. Excessive unexcused absences may affect your participation grade and overall course performance. Active participation in discussions, lab partnerships, and office hours is valued and contributes to the quizzes/participation component.
Late Work and Extension Policy
Assignments are due at the date and time listed on the schedule. Late submissions are accepted up to 48 hours past the deadline with a 10% penalty per day (i.e., 10% off for submission 0–24 hours late; additional 10% for submission 24–48 hours late). Work submitted more than 48 hours late receives no credit unless prior arrangements have been made with the instructor. If you anticipate difficulty meeting a deadline, contact the instructor before the deadline passes to negotiate an extension.
Capstone milestones (proposal, progress report, final report and presentation) follow the same late-work policy. Missing a midterm or capstone milestone without prior notice or valid excuse may result in a zero or significant grade reduction.
Regrade Request Policy
If you believe an assignment or exam has been graded in error, submit a regrade request to the instructor within 5 business days of receiving your grade. The request must clearly explain the basis for the appeal (e.g., incorrect scoring, misapplied rubric, factual error). Upon review, the instructor may adjust the grade upward, downward, or leave it unchanged. Regrade requests submitted after the 5-day window will not be considered.
Exam Format and Makeup Policy
This course does not include a traditional final exam. Assessment is based on homework, labs, quizzes, the midterm project, and the individual capstone project. If you must miss a quiz or capstone presentation due to illness, injury, or other documented emergency, contact the instructor immediately to arrange a makeup date or alternative submission format. Makeup assessments must be completed within 5 business days of the original due date unless otherwise agreed.
Academic Integrity
All work submitted in this course must be your own. Collaboration is encouraged where explicitly permitted (e.g., discussing concepts in office hours, peer code review during lab sessions), but presenting another person’s work, ideas, or code as your own—including unauthorized collaboration, copying, or reuse of solutions from online repositories, peers, or prior semesters—is a violation of academic integrity. Suspected violations are handled under the university’s academic integrity policy and may result in a failing grade on the assignment or in the course.
Generative-AI Use Policy
Generative-AI tools (such as ChatGPT, Claude, or GitHub Copilot) may be used only as explicitly permitted for a given assignment. When permitted, you must disclose which tool you used, how you used it (e.g., for brainstorming, code generation, debugging), and you remain fully responsible for the correctness and originality of everything you submit. Using these tools where they are prohibited, or submitting their output as your own unaided work, is a violation of academic integrity. The instructor will specify permissible AI use in the assignment brief for each homework and project.
Code Submission Standards
All programming assignments and the capstone project must adhere to the following standards:
- Version control: Code must be submitted via GitHub (or another instructor-approved platform) with a clear commit history. Include a README.md file documenting setup, dependencies, and instructions to run the code.
- Language and framework: Use Python 3.8 or later with TensorFlow 2.x or PyTorch 1.x (or instructor-approved alternatives).
- Documentation: Include inline comments and docstrings explaining key functions and design choices. Jupyter notebooks must be reproducible and include markdown cells describing the experiment or analysis.
- Reproducibility: All code must run without manual modification. Include a
requirements.txtorenvironment.ymlfile listing exact versions of dependencies. Provide random seeds and clear instructions for reproducing results. - Testing: Where appropriate, include unit tests or validation checks demonstrating that your code works as intended.
Failure to meet these standards may result in a grade penalty or rejection of the submission.
Professionalism and Conduct Expectations
Treat all instructors, peers, and guests with respect. Disruptive behavior, harassment, or discrimination of any kind will not be tolerated and may result in removal from class and referral to Student Conduct. In discussions and collaborative work, listen actively, assume good intent, and communicate professionally. When asking for help or providing feedback, be specific and constructive.
Communication Norms
Email response time: The instructor aims to respond to emails within 2 business days. For urgent matters, contact the instructor during office hours or before/after class.
Preferred channel: Use email for formal requests (extensions, makeup exams, regrade appeals) and Canvas messaging for quick clarifications about assignments. Office hours (time TBA) are available for in-depth discussions about course concepts and project guidance.
Disability Support Services
Students with disabilities may be eligible for academic accommodations. To request accommodations, contact the university’s Disability Support Services office or the equivalent accessibility services on your campus. Provide the instructor with any accommodation letter or documentation as soon as possible. The instructor will work with you to arrange reasonable accommodations such as extended exam time, note-taking support, or alternative assignment formats that do not compromise learning objectives.
Religious Holiday Accommodations
If a religious holiday or observance conflicts with a course deadline or assessment, contact the instructor at least two weeks in advance. Reasonable accommodations will be made to allow you to participate fully in the course without sacrificing your religious practice.
Counseling and Psychological Services
The university provides free, confidential counseling and mental health support for students experiencing stress, anxiety, depression, or other challenges. Information is available through the university’s health and wellness website. If you are struggling, please reach out to Counseling Services or speak with the instructor, who can direct you to appropriate resources.
Nondiscrimination and Harassment-Free Environment
The university is committed to a nondiscriminatory educational environment free from harassment, bullying, and discrimination based on race, color, religion, sex, national origin, age, disability, sexual orientation, gender identity, or other protected status. Any reported incidents will be addressed promptly and confidentially. If you experience or witness discrimination or harassment, report it to the Office of Equity and Inclusion or to the instructor.
Classroom Recording and Electronic Course Materials
Class sessions may be recorded by the instructor for asynchronous access and to aid students with documented disabilities. These recordings are intended for educational use within the course only and may not be shared, reposted, or distributed without explicit permission. Students are not permitted to record class sessions without prior written consent from the instructor. Course materials (syllabi, slides, handouts) are the intellectual property of the instructor and the university; unauthorized reproduction or distribution is prohibited.
Schedule
| Week | Session | Topic | Reading | Assignments / Due |
|---|---|---|---|---|
| 1 | Mon | Introduction and the Three Pillars of AI and Security | — | |
| 1 | Wed | Motivation, Course Overview, ML Refresher | — | HW1 out |
| 2 | Mon | Pillar 1: AI for Security — Foundations of ML for Cybersecurity | — | |
| 2 | Wed | Security Data, Feature Engineering, and Evaluation under Class Imbalance | — | HW1 due; HW2 out |
| 3 | Mon | Network Intrusion and Anomaly Detection | — | |
| 3 | Wed | Supervised vs. Unsupervised NIDS (Isolation Forest, Autoencoders, One-Class SVM) | — | HW2 due; HW3 out |
| 4 | Mon | Malware Detection and Classification | — | |
| 4 | Wed | Static/Dynamic Features, Deep Learning on Binaries (MalConv/EMBER) | — | HW3 due; HW4 out |
| 5 | Mon | Phishing/Fraud Detection (NLP, URL & Transaction Features) | — | |
| 5 | Wed | AI for Security Operations (Log Analysis, Alert Triage, Threat Intel, LLM Assistants) | — | HW4 due; HW5 out |
| 6 | Mon | Pillar 2: Security of AI — Threat Modeling for ML Systems | — | |
| 6 | Wed | Adversarial Examples: FGSM and PGD (attacking the Week 2–5 detectors) | — | HW5 due; HW6 out; Midterm Project Proposal out |
| 7 | Mon | Advanced Evasion: C&W, Transferability, Black-box Attacks | — | |
| 7 | Wed | Data Poisoning and Backdoor/Trojan Attacks | — | HW6 due; Midterm work session |
| 8 | Mon | Midterm Project Work Session | — | |
| 8 | Wed | Midterm Project Presentations | — | Midterm Project Report & Presentation due |
| 9 | Mon | Model Extraction, Inversion, and Membership Inference | — | HW7 out |
| 9 | Wed | LLM Security: Prompt Injection, Jailbreaks, Training-Data Extraction | — | |
| 10 | Mon | Adversarial Defenses: Adversarial Training and the Robustness Trade-off | — | |
| 10 | Wed | Certified Defenses, Adaptive Evaluation, and Poisoning/Backdoor Defenses | — | HW7 due; HW8 out |
| 11 | Mon | Pillar 3: Trustworthy AI — Interpretability: Intrinsic vs. Post-hoc | — | |
| 11 | Wed | LIME, SHAP, Attention, and Feature Attribution (Integrated Gradients, Grad-CAM) | — | HW8 due; Capstone Project Proposal out |
| 12 | Mon | Fairness: Definitions and Metrics (Parity, Equalized Odds, Calibration) | — | |
| 12 | Wed | Bias Detection and Fairness Audits; Impossibility Results | — | Capstone Project Proposal due; Capstone Work Begins |
| 13 | Mon | Fairness Mitigation and Bias Correction | — | |
| 13 | Wed | Reweighting, Adversarial Debiasing, Threshold Adjustment, and Trade-offs | — | |
| 14 | Mon | Privacy-Preserving ML: Differential Privacy and DP-SGD | — | HW9 out |
| 14 | Wed | Federated Learning, Secure Aggregation, Homomorphic Encryption | — | Capstone Progress Report due |
| 15 | Mon | Robustness Certification and Formal Verification (SMT, IBP, Randomized Smoothing) | — | |
| 15 | Wed | Neural Network Verification, Safety, and Certification | — | HW9 due |
| 16 | Mon | Transparency and Accountability (Model Cards, Audits, EU AI Act / NIST AI RMF) | — | |
| 16 | Wed | Three-Pillar Synthesis and Capstone Integration | — | |
| 17 | Mon | Capstone Presentations and Q&A | — | Capstone Report due; Capstone Presentation |
| 17 | Wed | Capstone Presentations and Closing | — |
Lecture Hints
Week 1: Introduction and the Three Pillars of AI and Security
Frame the whole course around three pillars: (1) AI for Security — using ML to defend systems (intrusion/malware/phishing detection, security operations); (2) Security of AI — protecting ML models from adversarial, poisoning, and privacy attacks; (3) Trustworthy AI — making AI safe, fair, robust, transparent, and accountable. Motivate with high-stakes deployments (SOC automation, autonomous vehicles, lending). Use one running example (e.g., an ML-based malware detector deployed in a SOC) that touches all three pillars: it must work (Pillar 1), resist evasion/poisoning (Pillar 2), and be explainable/accountable (Pillar 3). Recap the ML background the course assumes and preview the assessment arc.
Week 2: Foundations of Machine Learning for Cybersecurity ( AI for Security )
Pillar 1 (AI for Security). Establish the security ML pipeline: data sources (network flows/PCAPs, binaries, logs, emails), labeling challenges, and feature engineering for security data. Stress evaluation under severe class imbalance — precision/recall, PR curves, ROC/AUC, the base-rate fallacy, and the operational cost of false positives in a SOC. Discuss concept drift and non-stationary/adversarial data. Set up the toolkit students reuse for HW2–HW4.
Week 3: Network Intrusion and Anomaly Detection ( AI for Security )
Pillar 1 (AI for Security). Signature vs anomaly-based detection. Supervised classifiers (trees/gradient boosting, MLPs) vs unsupervised anomaly detection (isolation forest, autoencoders, one-class SVM) on network-flow data (e.g., NSL-KDD / CIC-IDS-style). Cover feature extraction from flows, handling rare attack classes, and setting operational thresholds / managing alert volume. Briefly preview that these detectors can themselves be evaded (foreshadow Pillar 2).
Week 4: Malware Detection and Classification ( AI for Security )
Pillar 1 (AI for Security). Malware detection and family classification. Static vs dynamic analysis features (byte n-grams, opcode sequences, API/system calls, PE headers, sandbox traces). Deep learning on raw bytes (MalConv-style) and the EMBER feature paradigm. Discuss packing/obfuscation, label noise, dataset bias, and temporal/eval pitfalls (train-test splits by time). Reiterate that the detector is an attack surface.
Week 5: Phishing/Fraud Detection and AI for Security Operations ( AI for Security )
Pillar 1 (AI for Security), closing the pillar. NLP and URL/feature models for phishing and spam; anomaly/fraud detection on transactions. Then AI for the SOC: log analysis, alert triage and prioritization, threat-intelligence extraction, and emerging LLM-based assistants for security operations — with caution about hallucination and prompt injection (bridges to Pillar 2). Summarize where ML genuinely helps defenders and its limits (evasion, drift, base rates).
Week 6: Threat Modeling for ML Systems and Adversarial Examples ( Security of AI )
Pivot to Pillar 2 (Security of AI). Threat-model ML systems: assets, adversary knowledge (white/grey/black-box), capabilities, and the attack surface across the ML lifecycle (data, training, deployment). Define adversarial examples and perturbation budgets (L_inf / L2). Implement FGSM (one-step) and PGD (iterative). Powerful framing: attack the AI-for-Security models students built in weeks 2–5 to show that defensive ML is itself a target.
Week 7: Advanced Evasion, Transferability, and Data Poisoning/Backdoors ( Security of AI )
Pillar 2 (Security of AI). Optimization-based attacks (C&W), transferability across models, and query-based black-box attacks. Then training-time threats: data poisoning (availability vs integrity) and backdoor/trojan attacks with triggers. Discuss real-world feasibility and ML supply-chain risk (pretrained weights, datasets). Implement a small backdoor on a toy model and observe trigger behavior — lab only, framed defensively.
Week 8: Midterm Project Work Session and Presentations ( Security of AI )
Midterm work session and presentations. Students present an ML security system they built (Pillar 1) and/or a threat model with an attack or defense (Pillar 2). Encourage peer questions; give feedback on threat-model rigor, attack strength, baseline soundness, and evaluation quality.
Week 9: Model Extraction, Inversion, Membership Inference, and LLM Security ( Security of AI )
Pillar 2 (Security of AI). Confidentiality attacks ON models: model extraction/stealing, model inversion (reconstructing training data), and membership inference (was this record in training?). Connect membership inference/inversion forward to privacy (Pillar 3, week 14). Then LLM-specific security: prompt injection (direct and indirect/RAG), jailbreaks, and training-data extraction. Discuss why deployed AI-for-Security tools that use LLMs inherit these risks.
Week 10: Adversarial Defenses and Robustness ( Security of AI )
Pillar 2 (Security of AI), closing the pillar. Defenses against evasion: adversarial training (PGD-based) and the robust-vs-clean accuracy trade-off; certified defenses (randomized smoothing, interval bound propagation) at a high level. Explain why obfuscated-gradient / gradient-masking defenses fail and why strong adaptive evaluation is essential. Defenses against poisoning/backdoors (data sanitization, spectral/trigger detection). Note deeper certification returns in week 15.
Week 11: Interpretability and Explainability ( Trustworthy AI )
Begin Pillar 3 (Trustworthy AI). Intrinsic interpretability (decision trees, linear models) vs post-hoc explainability (model-agnostic). Walk through LIME (local linear approximation) and SHAP (Shapley values); deep-net attribution (saliency, Integrated Gradients, Grad-CAM); attention as explanation and its limits. Stress faithfulness vs plausibility, and how explanations can expose bias and vulnerabilities — connecting back to Pillars 1–2 (e.g., explaining a malware detector’s decisions).
Week 12: Fairness: Definitions, Metrics, and Detection ( Trustworthy AI )
Pillar 3 (Trustworthy AI). Define group fairness rigorously: demographic/statistical parity, equalized odds, equal opportunity, and calibration within groups; group vs individual fairness. Explain impossibility results and fairness–accuracy / fairness–fairness trade-offs. Detect bias via stratified/subgroup evaluation with confidence intervals; discuss measurement bias and proxy variables. Use case studies (hiring, lending), and note fairness concerns in security ML (e.g., biased fraud or abuse detection).
Week 13: Fairness: Mitigation and Bias Correction ( Trustworthy AI )
Pillar 3 (Trustworthy AI). Mitigation across three families: pre-processing (reweighting, resampling), in-processing (constrained optimization, adversarial debiasing), and post-processing (threshold adjustment). Implement at least one and measure the effect on fairness metrics AND accuracy, framed as a Pareto trade-off requiring stakeholder input. Emphasize that mitigation may not be Pareto-optimal and that the right operating point is a sociotechnical decision.
Week 14: Privacy-Preserving Machine Learning (Differential Privacy, Federated Learning) ( Trustworthy AI )
Pillar 3 (Trustworthy AI). Motivate with regulation (GDPR, CCPA) and the model attacks from week 9 (membership inference, model inversion). Differential privacy formally ((epsilon, delta)); DP-SGD (gradient clipping + Gaussian noise) and the privacy–utility trade-off / privacy budget. Federated learning, secure aggregation, and homomorphic encryption at a high level. Emphasize privacy is a provable, composable property — unlike ad-hoc anonymization. Note the privacy epsilon differs from the adversarial perturbation epsilon.
Week 15: Robustness Certification, Formal Verification, and Safety ( Trustworthy AI )
Pillar 3 (Trustworthy AI). Verify neural-network properties: SMT solvers / symbolic methods, abstract interpretation and interval bound propagation for sound over-approximation, and certified robustness (randomized smoothing, the ACAS Xu benchmark). Discuss soundness vs completeness and the scalability limits of formal methods on large networks. Connect certification to safety cases and regulatory compliance (assurance, aviation, the EU AI Act).
Week 16: Transparency, Accountability, and AI Governance; Three-Pillar Synthesis ( Trustworthy AI )
Pillar 3 (Trustworthy AI) and synthesis. Transparency and documentation (model cards, datasheets, audit trails), accountability and auditing, and the regulatory/standards landscape (EU AI Act, NIST AI RMF). Synthesize all three pillars: AI-for-Security systems (Pillar 1) must themselves be secured (Pillar 2) and made trustworthy (Pillar 3). Guide capstone integration across pillars and provide actionable feedback on technical depth and communication.
Week 17: Capstone Presentations and Closing
Capstone presentations (10–15 min + Q&A) and course closing. Ask probing questions about threat models, design trade-offs, limitations, and real-world deployment. Recap the three pillars and how they interlock, then point students to where to go next: research venues, open-source tools, and communities. Provide actionable feedback on technical depth and presentation clarity.
End of Syllabus