top of page

GPT-4 Lung Biopsy Planning Found Shorter Needle Routes, but Two Failed Review

4 hours ago
10 min read

GPT-4 lung biopsy planning produced acceptable proposed needle paths in 28 of 30 cases, but the two failures exposed a consequential safety gap. The model worked from annotated, two-dimensional CT images rather than controlling a needle or evaluating patients during live procedures.

Researchers at Memorial Sloan Kettering Cancer Center retrospectively tested whether GPT-4 could recommend entry angles for CT-guided lung biopsies. Independent radiologists judged 93.3 percent of its proposed paths adequate for performing a biopsy.

The AI routes crossed a median of 5 millimeters of lung tissue, compared with 23 millimeters for routes previously selected by experienced clinicians. However, shorter does not automatically mean safer. One rejected route crossed four times more lung tissue than the corresponding manual route, while another created a potential needle-stability problem.

That tension matters more than the headline percentage. The study suggests a general-purpose multimodal model can interpret a carefully prepared medical image and produce a plausible geometric plan. It does not show that GPT-4 improves outcomes, reduces complications, or can plan biopsies without extensive human preparation and supervision.

What the GPT-4 Lung Biopsy Planning Study Actually Tested

The experiment evaluated proposed entry angles on historical images, not AI-guided procedures on patients.

The open-access study appeared in CVIR Oncology and involved 30 consecutive lung biopsy procedures. Researchers used de-identified images from procedures already performed by three interventional radiologists.

Each clinician had at least ten years of experience. The patient cohort included 17 men and 13 women, with a median age of 65 years. Lung nodules ranged from 5 to 45 millimeters in diameter, with a median diameter of 21 millimeters.

The sample included five apical lesions, seven subpleural lesions, three para-mediastinal lesions, and four lesions near the lung base. These locations introduced different planning challenges involving ribs, vessels, lung tissue, and needle stability.

A CT-guided lung biopsy requires a physician to advance a needle through the chest wall and into a suspicious lesion. The route must reach tissue suitable for diagnosis while avoiding bones, pulmonary fissures, the mediastinum, and major blood vessels.

Planning also considers how much normal lung tissue the needle crosses. Longer intrapulmonary routes can expose more tissue, although route length represents only one element of procedural risk.

The study did not give GPT-4 an unmodified CT examination. A researcher selected an axial image, cropped and magnified it, and added a radial grid using GIMP. The image was converted from its original medical format into a PNG file.

One author then marked the target with a blue circle and outlined sensitive structures in red. These annotations gave the model information that a deployable system would eventually need to identify automatically.

The researchers supplied five planning criteria through a standardized prompt. The requested path had to reach the target, reduce travel through lung tissue, and avoid bones, fissures, and the mediastinum.

Each case started in a fresh chat session. Researchers could continue prompting when the first response failed to address every condition. The model required a median of two iterations, with an interquartile range of one to three.

The team recorded the first response that incorporated all requested parameters. Two experienced interventional radiologists, neither involved in the original procedures, independently reviewed the resulting paths while blinded to their origin.

This setup separates feasibility from clinical performance. The model proposed angles after human experts had chosen the image, marked the target, identified hazards, and formulated the rules. A radiologist remained necessary at every important stage.

The study therefore tested a narrow question: can a multimodal general-purpose model generate a plausible entry angle from a prepared image? It did not test autonomous segmentation, three-dimensional planning, live navigation, or robotic needle control.

That distinction prevents a promising planning experiment from becoming an unsupported claim about automated surgery.

The Shorter Routes Created the Study’s Strongest Signal

GPT-4 proposed much shorter paths through lung tissue, but the study did not establish that those paths produced safer clinical outcomes.

Independent reviewers rated 28 of the model’s 30 proposed routes as adequate for performing the biopsy. That result produced the study’s 93.3 percent feasibility figure.

The AI-generated paths crossed a median of 5 millimeters of lung tissue. The interquartile range extended from 0 to 25 millimeters.

By comparison, the manual paths used in the completed procedures crossed a median of 23 millimeters. Their interquartile range was 5 to 57 millimeters.

The difference was statistically significant, with a reported p-value of 0.0063. Paths with no lung-tissue traversal involved subpleural nodules, which sit next to the pleural surface.

A direct path into such a lesion can avoid crossing normal lung parenchyma, the functional tissue involved in gas exchange. That makes a short route geometrically attractive, but geometry alone cannot define the best clinical approach.

The mean AI entry angle was 179 degrees, compared with 174 degrees for the routes clinicians used. That difference was not statistically significant.

Similar averages did not mean the model reproduced human decisions. Only eight of the 30 cases showed good concordance under the study’s stated comparison, according to the main results.

The authors defined a match using entry angles within ten degrees and a similar route through tissue planes. Elsewhere, they reported an even stricter match interpretation of three cases, highlighting sensitivity to the chosen definition.

This discrepancy deserves attention because central averages can conceal substantial case-level differences. Two groups of routes can have similar mean angles while diverging widely for individual patients.

The AI therefore appeared to find alternative routes, not merely imitate the completed procedures. Some alternatives looked shorter and remained acceptable to reviewers. Others exposed weaknesses that the aggregate result could easily hide.

The comparison was also structurally unequal. The manual routes had been executed successfully on patients, while the AI routes existed only as lines drawn over planning images.

A proposed route can appear reasonable in one axial slice yet become impractical when assessed across adjacent slices. Real procedures must account for respiratory movement, body position, needle anchoring, equipment constraints, and changing anatomy.

The original procedures produced no major complications. Four patients developed minimal pneumothorax that resolved without intervention, and two experienced mild, self-limiting hemoptysis.

Those outcomes cannot be attributed to the AI because the model did not guide the procedures. The researchers also did not simulate whether using its routes would have changed those complications.

The shorter-path result is consequently a useful hypothesis. It supports testing whether automated planning can reveal routes that clinicians might overlook, particularly when several geometrically valid options exist.

It does not support saying GPT-4 made lung biopsies safer. Establishing that claim requires prospective studies that compare clinical outcomes under controlled conditions.

Two Rejected Paths Show Why Shortest Is Not Always Safest

The model’s failures reveal a conflict between optimizing a visible metric and understanding how a needle behaves during an actual biopsy.

In Case 19, GPT-4 suggested a path that crossed 20 millimeters of lung tissue. The manual route crossed only 5 millimeters.

The AI recommendation therefore traveled four times farther through lung tissue, contradicting the optimization goal supplied in the prompt. It also returned a broad range of possible entry angles rather than a precise, clearly preferred plan.

Reviewers considered the route theoretically feasible but not representative of good clinical practice. That distinction captures the gap between avoiding obvious obstacles and producing a dependable procedural plan.

Case 27 exposed a different problem. The target was a subpleural nodule, meaning it sat close to the lung’s outer surface.

The proposed AI route unnecessarily crossed 5 millimeters of lung tissue. Reviewers also raised concerns about needle stability.

A very short route through supporting tissue can leave a coaxial biopsy needle insufficiently anchored. The needle can become more vulnerable to movement or dislodgement during sampling.

This creates the study’s central tradeoff. Reducing the distance through normal lung tissue sounds inherently beneficial, but a stable route can require enough tissue engagement to hold the needle securely.

The model received simplified rules rather than a complete representation of procedural practice. It could search for an angle that satisfied those rules without understanding every reason a radiologist might choose a less direct approach.

The researchers acknowledged that the analysis remained two-dimensional. Human radiologists normally scroll through multiple CT slices and use multiplanar reconstructions to understand the anatomy in three dimensions.

A single axial image cannot fully represent structures extending above or below that plane. It also cannot show how an apparently clear line changes across the needle’s complete three-dimensional course.

The images were manually annotated before reaching GPT-4. The model did not independently locate every lesion, segment the organs, or reliably define safety margins around critical anatomy.

The study also omitted a specified minimum five-millimeter margin from the subclavian and axillary neurovascular bundles. Patient positioning kept those structures away from the evaluated routes, but future systems cannot assume the same condition.

Prompting introduced another source of variability. Researchers began with a standardized prompt, but follow-up interactions were not fully scripted when the model missed a criterion.

Two users could therefore steer the same model toward different answers. A clinical planning system would need a reproducible process that produces auditable results without improvised conversation.

The cases were processed in separate sessions, which reduced cross-case contamination. However, the study did not establish whether repeated runs on the same image would return the same angle.

Consistency matters when a recommendation affects an invasive procedure. A system that changes its plan between runs needs a method for explaining, ranking, and resolving those differences.

The model version presents another challenge. General-purpose AI services change over time, sometimes without exposing every technical modification to clinical users.

A result obtained from one GPT-4 configuration cannot automatically transfer to a later model. Medical validation must attach to a controlled system version, defined inputs, and a documented operating process.

Purpose-Built Systems Put the General Model in Perspective

GPT-4’s result is notable because it used a general model, but specialized algorithms already offer more structured and reproducible planning methods.

A 2024 path-planning study evaluated software built specifically for transthoracic lung biopsy. It combined three-dimensional lesion detection with Bayesian optimization for proposing trajectories.

That system was trained using 2,147 nodules from 219 scans. Researchers validated lesion detection with another 235 scans containing 354 lesions.

The model achieved a reported area under the receiver operating characteristic curve of 97.4 percent. Its mean sensitivity was 93.5 percent, while mean specificity reached 93.2 percent.

Researchers compared proposed trajectories with actual biopsy routes from 150 patients. They classified a match as an angular difference below five degrees.

The system generated feasible routes in 85.3 percent of evaluated cases. Eighty-two percent matched actual paths under the study’s definition, with a mean angular deviation of 2.30 degrees among matching routes.

That system addressed tasks GPT-4 did not perform, including three-dimensional segmentation and lesion detection. However, it also remained a retrospective evaluation rather than proof of improved patient outcomes.

Earlier work likewise tested rule-based AI planning against specialist judgment. A 2023 proof-of-concept study generated five candidate routes for each of 28 patients.

Four experienced physicians rated 140 paths. The computer and physicians gave identical scores in 57.9 percent of cases, and all 28 algorithm-selected best paths were considered safe.

More recent research has expanded that comparison. The Trajectory Recommendation Algorithm for CT-guided Biopsy, known as TRAX, generates and ranks about 20,000 candidate routes after a physician identifies the target.

In a 53-patient TRAX comparison, blinded radiologists rated every AI-selected and physician-selected route as safe. Ratings were identical in 58 percent of comparisons.

Reviewers preferred TRAX in 23 percent and the physician route in 19 percent. The difference was not statistically significant, but TRAX routes were slightly shorter on average.

TRAX also selected more paths outside the standard axial plane. Half of its routes used an angled plane, while 81.1 percent of physician routes stayed axial.

These studies do not produce a simple contest with one winner. Their inputs, definitions, datasets, and route-matching thresholds differ.

A purpose-built system can encode anatomical constraints, quantify distances, and rank thousands of routes consistently. A multimodal language model offers greater flexibility but depends more heavily on image preparation, prompting, and human interpretation.

GPT-4’s role may therefore be closer to a planning assistant than a trajectory engine. It could help clinicians compare options, identify overlooked approaches, or document the reasons behind a selection.

Even that supporting role requires safeguards. The interface must distinguish model suggestions from approved plans, display the complete anatomical route, and preserve a record of human review.

The meaningful competition is not AI against radiologists. It is automated geometric suggestions against the complete clinical judgment required to turn a line on an image into a safe procedure.

GPT-4 Lung Biopsy Planning Still Needs Prospective Validation

The next evidence must show repeatability, three-dimensional performance, and patient benefit rather than another high feasibility percentage.

The first signal to watch is reproducibility. Researchers should rerun identical cases across controlled sessions and measure how often the model selects the same route.

They should also standardize every prompt and follow-up condition. If human correction remains necessary, studies must record how often it occurs and how much it changes the recommendation.

A reproducible system would strengthen the case for using generative AI within a formal planning workflow. Large variation between runs would weaken it, even if reviewers considered most individual outputs plausible.

The second signal is native three-dimensional analysis. Future systems need access to full CT volumes rather than manually selected PNG images.

That workflow should automatically segment the lesion, lungs, vessels, ribs, fissures, mediastinum, and other relevant structures. It should also calculate explicit safety margins and display the entire route.

Combining a language model with dedicated medical segmentation could address some of the current study’s artificial preparation. It would also create new integration risks that require separate validation.

Three-dimensional analysis must handle different scanners, reconstruction settings, patient positions, and lesion locations. Multicenter data would be essential because performance at one cancer center may not transfer to other hospitals.

The third signal is a prospective clinical comparison. Such a study would assess AI-assisted plans before procedures, while keeping qualified radiologists responsible for every final decision.

Researchers could compare diagnostic yield, pneumothorax, bleeding, procedure duration, radiation exposure, needle repositioning, and recovery time. They should report unsuccessful recommendations separately rather than letting averages obscure them.

Prospective testing should begin cautiously. A shadow mode could generate recommendations without influencing care, allowing investigators to compare AI plans with real decisions under live conditions.

Later trials could evaluate whether clinicians using the system make better or more consistent plans. Any movement toward direct robotic execution would require substantially stronger evidence and device-level oversight.

The 30-case study offers no answer about diagnostic yield because the proposed paths were not used. It also cannot show whether a shorter route would have reduced the observed minor complications.

Its value lies in establishing that a general multimodal model can participate in a tightly constrained planning exercise. That finding justifies better experiments, not immediate clinical deployment.

The Real Advance Is a Testable Planning Hypothesis

This feasibility study turns generative AI trajectory planning into a measurable research question while leaving clinical responsibility firmly with specialists.

GPT-4 identified routes that independent experts considered adequate in 28 of 30 historical cases. Its median proposed route crossed 18 millimeters less lung tissue than the completed manual routes.

Those figures make the method worth investigating. They also sit beside two rejected plans, manual image preparation, unscripted follow-up prompts, and a two-dimensional view of three-dimensional anatomy.

The most useful interpretation is neither dismissal nor celebration. A general-purpose AI model recognized enough visual and procedural structure to suggest plausible alternatives, but it lacked the complete context needed for dependable planning.

For radiologists, the study points toward decision support that compares trajectories rather than replacing clinical judgment. For medical AI teams, it defines the engineering work still missing between a chatbot experiment and validated software.

For hospitals and patients, the standard should remain outcome-based. Does AI assistance improve tissue sampling, reduce complications, shorten procedures, or make expert planning more consistent across care settings?

The next studies should answer those questions with controlled systems, full CT volumes, external datasets, and prospective clinical evaluation. Until then, GPT-4 lung biopsy planning remains an intriguing second opinion on a prepared image, not a navigator for a live needle.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page