Blog · Pen Testing

    Fuzz Harness Generation for Medical Device Protocols

    How to build FDA-defensible fuzz harnesses for the protocols medical devices actually speak.

    Abstract digital network connections and a stylized goat head illustrate fuzz harness generation for medical device cybersecurity
    On this page
    Christian Espinosa, Founder & CEO at Blue Goat Cyber

    By Christian Espinosa, MBA

    Founder & CEO · Blue Goat Cyber

    Published: · Updated:

    Key Takeaways

    • Fuzz testing is part of the security-testing expectations under the Feb 2026 guidance, scoped to the threat model and documented to the same evidence bar as the rest of security testing.
    • Generic fuzzers fail on medical protocols when they cannot handle statefulness, framing, checksums, association negotiation, MTU constraints, and block-wise transfer.
    • The defensible model is a harness per protocol, with explicit grammar sources, seed corpora, transport setup, and coverage signals.
    • AI helps with grammar synthesis from pcaps and specs, seed expansion, and crash deduplication. It does not reliably handle stateful orchestration or hardware-in-the-loop control.
    • Reviewers need harness documentation, seed corpus statistics, coverage reports, crash logs, and fixed-versus-accepted disposition tied back to the [threat model](/services/medical-device-threat-modeling "FDA-aligned threat modeling services") and security risk file.
    • The deliverable belongs in your SPDF and feeds your CAPA and VEX workflows after launch.
    Direct Answer

    A defensible fuzz harness establishes the session, preserves message framing, delivers mutations to the intended parser, and records how the device responds. We build it around the device’s actual protocols, such as HL7, DICOM, BLE GATT, MQTT, or proprietary serial links, using specifications and captured traffic. The evidence must show coverage, reproducible findings, and disposition traced to threat-model entries and security risk controls.

    Published June 3, 2026

    Why this matters

    A fuzzer can send traffic all night without reaching the code we intended to test. If the device rejects every message at authentication, framing, or session negotiation, the execution count tells us little about parser coverage. The harness is what gets the test past those gates and makes the result repeatable.

    The FDA’s Cybersecurity in Medical Devices: Quality Management System Considerations and Content of Premarket Submissions (Feb 3, 2026 final guidance) places security testing within the cybersecurity evidence expected for applicable submissions. Section 524B of the FD&C Act establishes statutory requirements for cyber devices. We treat fuzzing artifacts as controlled engineering evidence, alongside the software lifecycle work under IEC 62304 and security risk-management work under AAMI TIR57 and ANSI/AAMI SW96:2023.

    Cybersecurity deficiencies contribute to first-cycle Additional Information (AI) requests. The FDA's FY2025 MDUFA Performance Report lists cybersecurity among the subject-matter expert areas that may be involved in reviewing a device submission, alongside sterility, biocompatibility and software. That broader pattern is not evidence that fuzz-harness gaps alone are the most common deficiency. For this work, the practical question is narrower: can a reviewer determine what the campaign exercised and what happened when it found a problem?

    What the FDA actually expects for fuzz testing evidence

    We document fuzz testing within the security-testing package supporting Section 524B and the Feb 3, 2026 final guidance, Cybersecurity in Medical Devices: Quality Management System Considerations and Content of Premarket Submissions. A tool name and a runtime are not enough. The report needs:

    • Scope tied to the threat model and architecture views: interfaces, protocols, trust boundaries, and any exclusions with their rationale.
    • Methodology: mutation, generation, coverage-guided, or protocol-aware testing; the harness; the seed corpus; and the runtime environment.
    • Duration and intensity: executions, total runtime, throughput, and coverage achieved.
    • Findings and disposition: crashes, hangs, and undefined behavior, each tracked to a fix, a compensating control, or documented risk acceptance.
    • Tester independence and qualifications: who performed the testing, their relationship to the design team, and who takes responsibility for the report.

    In the FDA letters we reviewed for our September 23, 2026 MTEC webinar, one letter raised inadequate control detail nine separate times against nine controls in a single submission. The recurring problem was a principle without a testable specification: passwords without a password policy, TLS 1.2 without cipher suites, or ECDSA without the curve and hash function. This was not a fuzz-specific finding, but the lesson applies directly. A harness cannot demonstrate that a control meets its requirement if the requirement never defines the expected behavior.

    “We ran AFL++ for an hour on the management port” leaves the same gap. We need to show which messages reached which states, what behavior counted as a failure, and how the result traces to a requirement or threat.

    Why generic fuzzers fail on medical protocols

    The problem is not that an off-the-shelf fuzzing engine is unusable. It is that the engine usually needs protocol-specific code around it. Medical devices speak negotiated, stateful, framed protocols over transports that reject malformed traffic before it reaches the intended target.

    • Stateful handshakes. DICOM requires Association establishment through A-ASSOCIATE-RQ/AC and presentation-context negotiation before P-DATA flows. HL7 v2 over MLLP needs start/end framing and an ACK/NAK loop. MQTT needs CONNECT/CONNACK and maintains session state. A stateless mutator may exercise only rejection paths.
    • Checksums and length prefixes. Proprietary binary protocols, BLE GATT extensions, and serial protocols can include CRCs, lengths, and sequence numbers. A payload mutation that invalidates the wrapper may never reach the payload parser.
    • MTU and fragmentation. BLE ATT starts with a 23-byte default MTU; larger values require negotiation, and characteristic-value limits also constrain transfers. CoAP uses Block1/Block2 for block-wise transfer when needed. The harness must respect the target’s supported boundaries or deliberately test violations while recording how they are handled.
    • Association and presentation contexts. DICOM needs the appropriate SOP Class UID and transfer syntax in an accepted presentation context. A C-STORE sent for an unaccepted context does not exercise the intended object parser.
    • TLS and authentication wrappers. MQTT and many HL7 v2 deployments use TLS, sometimes with mutual authentication. Without the required credentials and session setup, the fuzzer tests the wrapper rather than the application protocol.
    • Hardware-in-the-loop state. BLE testing needs a controlled radio and a known pairing, bonding, power, and connection state. CAN/CANopen needs bus access and termination. Recovery from a reset must restore those conditions before the next iteration.

    In our representative Class II wearable testing, the device used BLE 5.2 for telemetry, configuration, and firmware updates through a mobile companion app. We used a rooted Android test phone, BTLEJuice, and custom GATT clients for active fuzzing, with BLE captures to observe the exchanges. That setup matters because generating characteristic values is only one part of the job. The test also needs to reach the intended characteristic in a known connection state and distinguish a rejected write from a device failure.

    A harness handles those prerequisites so the fuzzing engine can spend its time on the selected parser or state machine.

    Harness generation per protocol

    We start with the protocol specification, the implementation’s supported features, and valid traffic. Then we choose the engine. The table below lists practical starting points, not interchangeable recipes.

    In the table, “use” refers to the harness. Its framing, state, transport, and failure detection are the engineering deliverable. Response codes and disconnect rates are useful observations, but we do not present them as equivalent to measured code coverage.

    ProtocolTransportStatefulnessGrammar sourceTooling startSeed corpusCoverage signal
    HL7 v2MLLP / TCP (often TLS)Framed, ACK/NAK loopv2.x message profiles, conformance statementboofuzz with MLLP wrapperADT, ORU, ORM, MDM exemplars from specACK rate, parser exceptions, log delta
    DICOMTCP (DUL), optional TLSAssociation negotiation, presentation contextsDICOM Part 5/7/8, IODs, SOP Classespynetdicom + boofuzz, custom mutatorValid C-STORE/C-FIND per accepted contextAssociation success rate, SCP error responses
    BLE ATT / GATTBLE link layer + L2CAPPaired/bonded, MTU-negotiatedService/characteristic enumeration, device specDefensics BLE, sweyntooth-style use, custom on nRF52/ESPValid characteristic writes, indication round-tripsDisconnect rate, link supervision timeouts, watchdog resets
    MQTTTCP, almost always TLSSession, QoS state, retained messagesMQTT 3.1.1 / 5.0 spec, broker schemaboofuzz or libFuzzer with TLS shimValid CONNECT, SUBSCRIBE, PUBLISH per topic treeCONNACK reasons, broker logs, ACL violations
    CoAPUDP, optional DTLSBlock-wise transfer, observeRFC 7252 + 7959, device resource mapAFL++ persistent on a CoAP library wrapperValid GET/POST per resource, multi-block PUTResponse code distribution, DTLS handshake errors
    Proprietary binarySerial, USB-CDC, sub-GHz RF, CANCustom framing, checksums, sequenceReverse-engineered from pcaps + firmwareCustom use over boofuzz blocks or libFuzzerCaptured exemplars, mutated within framing constraintsCRC accept rate, command ACKs, controller reboots
    The fuzzer is the engine. The use - framing, state, transport, oracle - is the deliverable.

    A few harness-design rules that apply across all of these

    • Make state explicit. Encode session establishment, authentication, reset, and recovery. Confirm the target’s state before counting an iteration as a test of the intended path.
    • Separate wrapper tests from payload tests. For deeper parser testing, mutate the payload and recompute checksums and lengths. Test invalid checksums and lengths separately rather than letting them consume the whole campaign.
    • Define a failure oracle. Specify how we detect incorrect behavior: crash logs, watchdog resets, unexpected configuration changes, MQTT disconnect reasons, or DICOM SCP error classes. “The connection closed” needs investigation, not an automatic vulnerability label.
    • Record enough to reproduce. Preserve the seed, mutation, message sequence, transport capture, firmware identity, and target state. A payload alone may not reproduce a state-dependent failure.
    • Use real hardware where the threat model requires it. An RTOS parser running on x86 can support fast testing and triage. It does not establish coverage of radio behavior, device timing, or hardware-dependent recovery.

    State handling deserves its own tests. In our representative Class II wearable RF testing, a proprietary protocol used a fixed 32-bit header and a sequence number that reset on power-cycle. Captured telemetry packets replayed successfully to the base station. This was a replay finding, not a fuzz-discovered crash. It illustrates why a harness must record power-cycle and session state: testing only one uninterrupted session can miss behavior across resets.

    Where AI legitimately helps (and where it doesn't)

    We use AI for drafts and sorting tasks that engineers can check. We do not let generated code define the coverage claim.

    • Grammar synthesis from pcaps and specs. An LLM can draft boofuzz blocks, an ASN.1 sketch, or a Kaitai struct from captures and specifications. We check field boundaries, optional fields, constraints, and implementation-specific behavior before using it.
    • Seed corpus expansion. AI can draft varied HL7 v2 messages, DICOM objects per IOD, MQTT topic trees, and CoAP resource sets. We verify that supposedly valid seeds are accepted by the intended implementation.
    • Crash deduplication and triage. Clustering by stack trace, register state, and message prefix can reduce repetitive review. A proposed root cause still needs reproduction and engineering analysis.
    • Mutator authoring. AI can draft custom AFL++ or libFuzzer mutators that preserve framing. We test the mutator itself before trusting its output.

    The work that remains human-owned is less convenient to automate:

    • Stateful orchestration. DICOM Association establishment, BLE pairing and bonding, MQTT session restoration, and CAN/CANopen NMT transitions need explicit state machines and recovery logic.
    • Hardware-in-the-loop control. Power cycling, reconnecting a radio, restoring a UART logger, and detecting an incomplete reboot require bench instrumentation and tested scripts.
    • Embedded crash analysis. Without source, symbols, or a reliable debugger, an STM32 watchdog reset may require trace-buffer analysis at the bench. A log-line explanation is not a root cause.
    • Threat-model-aligned scope. An AI tool cannot decide which characteristics carry safety-relevant writes without device-specific evidence. We own that decision.

    The pattern is human-led and AI-assisted, with a named tester responsible for scope, coverage, and findings. Our companion piece on where AI penetration testing fails an FDA reviewer covers that boundary in more detail.

    Process flow: spec to VEX

    We keep the harness, corpus, campaign configuration, and report linked to the tested device build. Otherwise, the work becomes difficult to reproduce after a firmware change.

    From spec or pcap to CAPA and VEX

    1. 1. Spec and pcap to grammar. AI may draft the message grammar from RFCs, conformance statements, and captures. An engineer checks framing, constraints, and supported features.
    2. 2. Harness authoring. We encode the state machine, transport, TLS/DTLS setup, checksum fix-ups, recovery behavior, and device-side failure oracle.
    3. 3. Seed corpus. We collect valid exemplars per IOD, topic tree, characteristic, or resource. AI may expand them; an engineer verifies, prunes, and labels the results.
    4. 4. Campaign execution. We run against the selected target, using hardware-in-the-loop where required, with crash capture, power-cycle automation, and coverage logging.
    5. 5. Triage and deduplication. AI may cluster signatures and suggest causes. We reproduce the behavior and assess reachability, exploitability, and impact.
    6. 6. Reproducer. We preserve the smallest reliable reproducer together with the required target state so developers can investigate and we can retest.
    7. 7. CAPA and VEX. Findings enter the security risk process and, where applicable, CAPA and post-launch VEX workflows. Evidence supports not_affected, affected, or fixed status; a clean campaign alone does not establish not_affected.
    8. 8. Report and signature. A named, qualified tester takes responsibility for the methodology, coverage, findings, limitations, and disposition recorded in the report.

    Flow: Spec / PCAP → Grammar → Harness → Corpus → Campaign → Triage → Reproducer → CAPA / VEX

    Deliverables a reviewer wants to see

    See also: BLE & RF Penetration Testing, What Is MedTech? MedTech vs Medical Device vs Life Sciences, and 8 FDA Cybersecurity Deficiencies, Ranked From Real Letters.

    For each protocol, we assemble evidence that explains both the test and its limits:

    • Harness source or a controlled reference in the DHF: enough detail to inspect the state machine, framing, and failure oracle.
    • Seed corpus statistics: counts, sources, generation methods, and mappings from exemplars to threat-model paths.
    • Coverage report: basic-block, branch, or protocol-state coverage, as available. State clearly when only external behavior is observable.
    • Crash log and deduplication summary: unique signatures, reproduction reliability, and reachability from exposed interfaces.
    • Disposition: fixed, mitigated through a compensating control, or accepted with rationale traced to the security risk file.
    • Tester qualifications: independence and relevant experience, consistent with the rest of the penetration test report.
    • Tooling and environment: tool versions, target image hash, hardware bench configuration, and AI tools and model versions if used.

    This is the part teams skip: explaining what the campaign did not reach. If authentication blocked a path, an instrumented build differed from production, or a radio interface was tested only through an emulated parser, state that limitation and explain how the remaining risk was addressed.

    These artifacts support a response to an AI request, Major Deficiency, or Hold letter. They do not guarantee that a reviewer will accept the scope. They let the reviewer evaluate it.

    Need help? Our team supports manufacturers with FDA cybersecurity submissions end-to-end. Explore our medical device cybersecurity services or book a discovery call.

    Generic fuzzer vs protocol-aware harness

    The difference is whether mutations reach the intended code and whether we can explain the result. The table describes the intended comparison; a protocol-aware harness still needs checks showing that session setup and parser reachability actually worked.

    AspectGeneric network fuzzerProtocol-aware harness
    Gets past authentication or session setupRarelyYes, by construction
    Understands message framing and checksumsNo, so packets are dropped earlyYes, so mutations reach the parser
    Code reached in a typical runConnection handling onlyParsing, state machine, and error paths
    Crash triageManual, often duplicatedDeduplicated against the input that caused it
    Evidence quality for a reviewerHard to defend coverageCoverage tied to protocol states and interfaces
    Setup costLowReal, but reusable across releases

    How Blue Goat Cyber runs this

    We scope protocol-aware fuzzing against the device’s threat model and architecture views. For each declared interface, we determine the relevant protocol paths, required target state, available instrumentation, and hardware needs. The tester owns those choices and documents exclusions rather than hiding them behind a tool’s feature list.

    We build the harnesses, run the security tests, reproduce findings, and retest remediations. Where AI helps with grammar drafting, seed expansion, or crash clustering, we disclose its use in the report. The final coverage claim rests on observed testing, not generated code or execution counts alone.

    If you are scoping a 510(k), De Novo, or PMA submission and need fuzz testing with reviewable evidence, book a strategy session.

    To see how we test for these weaknesses on real hardware, see our medical device penetration testing service.

    FAQ

    Does the FDA require fuzz testing for medical devices?

    The February 3, 2026 final premarket cybersecurity guidance identifies fuzz testing within the expected security-testing approach. Section 524B establishes cybersecurity obligations for applicable cyber devices; it does not prescribe a particular fuzzer or campaign duration.

    We scope fuzz testing to the threat model and document the methodology, coverage, findings, and tester qualifications. For Class II and III connected devices, leaving relevant interfaces untested or omitting the rationale can lead to AI Requests or Major Deficiencies. The submission needs evidence for the chosen scope, not just a statement that fuzzing occurred.

    How is medical-device fuzzing different from IT or web fuzzing?

    Medical devices commonly use negotiated, framed, stateful protocols: DICOM, HL7 v2 over MLLP, BLE GATT, MQTT, CoAP, proprietary serial, and CAN/CANopen. They may run on embedded targets without accessible crash dumps or symbols and communicate over radios or physical buses.

    IT and web testing often has richer instrumentation around HTTP parsers, though those systems can also be stateful. In medical-device work, we usually spend more effort on transport access, state restoration, and observing failures. The fuzzing engine is only one component.

    Premarket vs postmarket fuzzing - what changes?

    Premarket fuzzing produces controlled evidence for the submission. Postmarket, we reuse the harnesses as regression assets when a third-party component receives a CVE, firmware changes, or a customer reports anomalous behavior.

    Keep the corpus and harness versioned with the device build. Reproducible tests can support VEX “affected” or “not affected” determinations, but absence of a crash does not by itself prove that a vulnerability cannot affect the device.

    How long should a fuzz campaign run per protocol?

    We use coverage progress and the planned test scope to set stopping criteria, not a universal clock. A mature parser may take tens of millions of executions to reach a plateau; a custom binary protocol may require days. Hardware recovery time can reduce throughput sharply.

    Document runtime and coverage together. “Branch coverage plateaued at X% after N hours, with no new crashes in M hours” is more useful than “we ran for a week.” If branch coverage is unavailable, identify the protocol states and observable behaviors used instead, and state that limitation.

    How do we tie fuzz findings to the security risk file?

    We map each unique finding to the relevant threat, vulnerability, patient-safety impact, control, and residual risk. This follows the AAMI TIR57 / ANSI/AAMI SW96:2023 approach.

    The record should distinguish the observed device effect from the potential patient harm. A reset is an observation; its safety significance depends on what the device was doing and how it recovered. If a finding has no corresponding threat-model entry, we update the model or document why the finding falls outside it.

    What does "fixed vs accepted" disposition look like for fuzz findings?

    We use three clear outcomes:

    • Fixed: the code or configuration changed, and security retesting confirms the original reproducer no longer causes the failure. We also rerun relevant corpus inputs.
    • Mitigated: a compensating control, such as an input-length cap, rate limit, or recovery mechanism, reduces the risk. We test that control and document the remaining risk.
    • Accepted: the responsible manufacturer personnel approve a rationale tied to the threat model, exploitability, and patient-safety impact.

    A watchdog reset is not automatically an acceptable mitigation. Its effect on the device’s clinical function still needs assessment. Reviewers need an evaluable disposition for every finding, not a report that leaves crashes untriaged.

    About the author

    Christian Espinosa, Founder & CEO at Blue Goat Cyber

    Christian Espinosa, MBA · Founder & CEO, Blue Goat Cyber

    U.S. Air Force Academy graduate and veteran with 30+ years in cybersecurity. Founded Alpine Security in 2014 (acquired 2020), then Blue Goat Cyber in 2022. Has supported 275+ medical devices, with no cybersecurity-related rejections to date. Author of three books including The Smartest Person in the Room. Ironman triathlete and mountaineer.

    Read more about ChristianLinkedIn

    Sources & references

    Primary sources cited in this article. Links open in a new tab.

    1. FY2025 MDUFA Performance Report- U.S. FDA

    Where your device stands

    Find your stage in the FDA cybersecurity journey

    Answer one question about where your device is today. You get your stage, the next action to take, and the support that fits it.

    Find where my device stands

    Got a deficiency letter?

    Get a free FDA cybersecurity deficiency letter review

    Paste the cybersecurity questions from your AI request, hold letter or refuse-to-accept notice. The triage tool sorts each question and outlines the evidence the reviewer is asking for. Want an expert to read it? We return a gap analysis within 48 hours.

    Related services

    Put this into practice on your device

    Every Blue Goat Cyber engagement maps directly to FDA Section 524B and the SPDF - so the evidence you need lands in your submission, not in a separate report.

    Ready when you are

    Get FDA cleared without the cybersecurity headaches.

    30-minute strategy session. No cost, no commitment - just answers from people who've shipped 275+ devices supported.