Introduction

It all started with the desire to get at least one CVE issued at a time when bug bounties were closing. I haphazardly downloaded an open-source project, asked Claude-Code to read its Security.md, and analyze/classify existing CVEs. Naturally, I received several vulnerability candidates similar to the existing CVEs.

However, during the actual implementation and verification process, I confirmed that they were not vulnerabilities for various reasons. After wasting tokens meaninglessly like this, I concluded that there were limitations to finding vulnerabilities through simple static source code analysis. I decided to create an open-source vulnerability analysis flow and use Local LLMs and MCP tools if necessary.

Coincidentally, among the open-source candidates I wanted to find vulnerabilities in, there was an AI flow open-source called Flowise. Initially, I thought a project to “find Flowise vulnerabilities using Flowise” would be interesting, so I proceeded with it. However, due to the performance of Local LLMs and the difficulty of debugging, I ended up pursuing additional learning and exploring other directions.


1) Test1 : Flowise & Local LLM(Qwen3:14b)

Settings

  • Flowise In docker container
  • Ollama
  • Local llm model Qwen3:14b
    • almost opensource llm model do not support tool calling or standard of MCP tool calling
    • While setting it up, I realized how amazing commercial LLM services are.
  • Burp Suite MCP
    • MCP bridge to Flowsie container (AI handmade, the given proxy and Burp Suite must follow a defined data format, but it didn’t match Qwen’s tool calling format, so I built a proxy)

Flowise

  • Custom MCP node ⇒ Tool agent Node’s Tools
    • The Custom MCP node calls Burp MCP through a custom-built bridge.
  • Ollama Node ⇒ Tool agent Node’s Tool calling chat model
    • Qwen3:14b, Context Window set to 16k tokens
  • Buffer Memory ⇒ Tool agent Node’s Memory

Test1 Challenging

  • Local LLM performance

    • Insufficient tokens made it difficult to grasp context.
    • Unlike commercial models, it wasn’t proficient in using MCP tools; models that support tool calling exist separately.
    • Meanwhile, I felt my MacBook, which rarely got hot, heating up.
  • Lack of LLM service design capability

    • Started with the need to reduce the use of commercial LLMs, but ultimately lacked overall learning on how to handle agents, which LLM model is suitable, etc.
    • Realized it was very difficult to harness the MCP tool part in an unprepared state.
    • Had no idea how to proceed with harnessing in Ollama.

⇒ Realized that I needed to set smaller goals and proceed. ⇒ It’s necessary to gradually establish a methodology for vulnerability analysis itself… but a lot of thought will be needed to run that part with local models. For Test2, I decided to try a simpler documentation task while studying AI and getting familiar with it.


2) Test2 : Change LLM Model(gemma4:e4b) & Specify the purpose to use Local LLM

Test2 Project Goal : Create AI vulnerability security review items

I had previously created AI security review items with NotebookLM, but the process of retrieving supporting clauses was unstable. I judged that relying entirely on Claude would also be incomplete, so I decided to let Claude handle the initial draft generation and final review, but use a Local LLM to generate RAG for accurate evidence retrieval and use the local LLM for simple, repetitive tasks.

Gemma4:e4b

  • Model memory limitations: My MacBook Pro uses VRAM and RAM together, totaling 24GB. Since it runs iOS and other tasks simultaneously, memory space is not abundant. If the model itself is heavy, context can be limited, and responses can be very slow or go off-topic.
  • While LLM performance is important, I believe harnessing is more crucial for local LLMs that already have performance limitations.
  • Speed: Aimed for fast speed by using e4b, designed for mobile/on-device use.

Thinking

Flow

  • Step1 (Claude) : PDF to Text python Script generation
> **참고자료**
> - `KISA` : 인공지능(AI) 보안안내서 (과기정통부·KISA, 2025.12)
> - `NIS` : AI보안 가이드북 (국가정보원, 2025.12)
> - `FSI-AG` : AI 에이전트 아키텍처에서 인증 및 권한관리를 위한 보안고려사항 (금융보안원, 2025.09)
> - `FSI-GL` : 금융분야 AI 보안 가이드라인 (금융보안원, 2023.04)
> - `FSC` : 금융분야 AI 개발·활용 안내서 (금융위원회, 2022.08)
> - `AI기본법` : 인공지능 발전과 신뢰 기반 조성 등에 관한 기본법 (2026.01.22)
> - `AI시행령` : 인공지능기본법 시행령 (2026.01.22)
  • Step2 (Claude) : WIKI structure generation
  • Step3 (Claude) : Prompt generation for Gemma e4b (summarizing files to fit Gemma’s context length)
  • Step4 (Gemma4) : WIKI creation (referencing reference text files)
  • Step5 (Claude): Draft md file for security review items, referencing WIKI data
  • Step6 (Claude , Gemma4) : Final markdown table review
    • bge-m3:latest (local) — RAG generation
      • RAG generation based on original text
    • gemma4:e4b (local) — Heavy repetitive tasks
      • Scoring md files based on WIKI and RAG (checking if they are in the actual report, if they are appropriate)
    • Claude Sonnet 4.6 — RAG design, task distribution to Gemma4:e4b, final report review and modification
      • Final markdown quality review (referencing scoring using RAG)

Scoring Result) These are the two items caught as partially compliant through scoring. It’s uncertain if the items were well-chosen, but the reasons are quite logical.

Test2 Results (Review and Vulnerability Check Checklist)


3) Find Vulnerabilities by Using AI

Find Vulnerabilities by Using AI

I wanted to organize what I learned from the previous two tests (1 and 2) and use it for test 3. While not groundbreaking, here’s what I actually felt:

  1. Use Local LLMs only for tasks that are very clear, simple, and repetitive.
  2. For tasks requiring thought and reasoning, always use commercial AI.
  3. MCP integration should be as easy as possible.
  4. RAG is good when semantic search and original source verification are needed.
  5. I understand the need for AI Agent Flow tools (Flowise, n8n, etc.), but simple calls are more convenient if debugging is important.
  6. HIL (Human In the Loop) is more necessary than I thought for me, who uses Claude Pro (Claude is dumber than I thought).
  7. It’s difficult to find meaningful vulnerabilities by simply searching based on existing patch logs.
  8. Caveman skills save a lot of tokens.
  • Initially, when analyzing an open-source project called Dify, I thought about building existing source code as RAG and searching with existing security diffs (Before Code), but this was a thought without a good understanding of AI. Contextual similarity does not mean code structure similarity.
  • I conducted inspections of open-source projects like Dify, Plane, n8n, and Langflow, and after some adjustments, the current FLOW is as follows.

Problem Definition

  • Target: Appealing open-source projects
  • Success Criteria: Reproducible PoC + Vulnerability report + CVE ID acquisition

Flow

For each Phase, I record what was done within the phase folder. This prevents duplication when sessions are interrupted or when searching for candidates.

  • Phase1. Environment Setup

    • Environment setup is done first.
    • Although my FLOW design skills are lacking, I judged that understanding what kind of service the open-source project is, is necessary first.
    • Set up with Docker / AWS, and if there are accounts, set up accounts with various permissions.
    • Briefly grasp the open-source structure and write a readme about it.
  • Phase2 Vulnerability Information Discovery

    • Existing vulnerability discovery
      • Use GitHub API to check GitHub security advisories and security commit logs.
      • Store vulnerabilities in a specific format ([CVE/GHSA ID] | [Type] | [Affected Module] | [Vulnerability Pattern] | [Fix Method]).
      • Count by type to prioritize.
    • Check areas likely to have vulnerabilities
      • New features
      • New files
      • Attack surface grep (file/upload/request, etc.)
  • Phase3 Information-based Vulnerability Candidate Location Discovery + In-depth Analysis

    • For existing vulnerabilities from Phase 2 (in priority order) + parts found in new files from Phase 2
      • Extract code characteristics from CVEs and grep
      • Semgrep
    • Claude performs in-depth analysis of vulnerability candidates found above.
    • Confirm where function input values come from.
    • Create final vulnerability candidates.
  • phase4 By design filter

    • While searching for vulnerabilities, there are cases where ~ is Out of Scope in Security.md, or the maintainer explicitly states “this is by design” for an issue raised.
    • Reviewed Security.md and designs related to the found vulnerability candidates.
  • phase5 Vulnerability POC

    • Plan and get approval to actually perform POC for the found vulnerability candidates.
    • Stated that if human help is needed during further progress, to ask.

3 vulnerabilities reported, 1 CVE, 1 DUP, 1 pending after fix

  • Langflow : Triage in progress, response and fix DRAFT in progress (6/4), CVE issued! (CVE-2026-12944)

  • n8n : Received a reply that it would take a long time because they were too busy… but it seems a similar vulnerability was patched later while fixing related code (6/20) As expected, received a “Dup” notification (07/13)

  • 50K project : Fix confirmed merged into main branch, received new GHSA number, not yet an issue (6/24)

Lessons learned from this project

  • When utilizing AI, harnessing is as important as performance.

    • Clearly define the purpose of using AI.
    • Fix the output and input formats required to achieve that purpose.
    • Identify AI bottlenecks and control prompts or memory (constantly waiting for unnecessary things).
      • It’s okay to take a slight detour, but if you can’t wait for work to proceed, you need to review whether it’s actually a task worth waiting for every 5 minutes using project md or a separate workflow, or proceed by logging.
    • If the above problem definition is clear + you have enough tokens (financial resources) or computing power to run multiple agents, then agentic engineering might be somewhat possible…
  • Problems arising from AI agentic engineering features

    • Although authentication/authorization was important, as server-to-server communication increases, the probability of vulnerabilities increases not only in client authentication but also in all end-to-end verification.
  • The ability to consider methodologies for AI risk assessment (practical inspection methods), or the ability to design controls for values resulting from LLMs during design (harnessing), is needed.

  • Open-source analysis using AI is possible, but I am skeptical about using AI to find vulnerabilities in internal systems as a person in charge in the future. There will be an enormous number of vulnerability candidates, and only the person reviewing them will suffer. This project also took about 30 hours to verify meaningful vulnerabilities out of 10.

  • It seems that abilities such as performing Triage using AI, software architecture skills to structurally eliminate vulnerabilities, or managing various authentication permissions in Cloud/Containers will become important.