AI-Generated Proposals Are Weakening Traditional Signals in Federal Acquisition
- Ethan Carter

- 5 days ago
- 13 min read
Google News surfaced a sharp warning about AI-enabled acquisition: better proposal language can now make genuine delivery capability harder to identify. The Federal News Network commentary frames this conflict as the “perfect proposal paradox” and connects it with “Garcia’s law.” The labels are provocative, but the underlying problem is practical. Contractors can produce polished, compliant responses faster, while agencies still must determine who can perform the work.
That changes the nature of a federal competition. Proposal quality once reflected several scarce capabilities, including institutional knowledge, disciplined planning, technical judgment, and the ability to coordinate experts under pressure. Generative AI can now imitate many outward signs of those capabilities. It can organize requirements, standardize terminology, rewrite weak passages, and identify missing sections within minutes.
The result is not simply an increase in automated writing. It is a collapse in the value of polish as a signal. When every bidder can produce clear prose and perfect formatting, evaluators must look elsewhere for evidence of understanding. The main contest is becoming AI-assisted presentation versus independently demonstrated performance.
Google News is only the discovery channel here, not the source of the argument. The important story concerns federal buyers, contractors, evaluators, and the public missions affected by their decisions. The government must gain AI’s speed without allowing persuasive text to become a substitute for evidence, accountability, or human judgment.
What the Google News Story Actually Changes
The perfect proposal is becoming easier to manufacture, so proposal perfection carries less information than it once did.
Generative AI already fits naturally into proposal work. A model can summarize a solicitation, extract requirements, draft compliance matrices, suggest themes, and compare sections against evaluation criteria. Retrieval systems can also ground outputs in a contractor’s approved project records, resumes, and past performance descriptions.
These functions reduce repetitive labor. They help teams find omissions before submission and give subject-matter experts a cleaner draft to review. Smaller contractors can also use them to compete with organizations that maintain large proposal departments.
The tension begins when assistance becomes substitution. A system can produce a credible technical narrative without knowing whether the proposed staff, schedule, architecture, or management controls will work. It predicts convincing language from available context. It does not accept contractual responsibility for the claims it generates.
That distinction matters because federal source selection is not a writing contest. Under FAR 15.305, proposal evaluation assesses both the submission and the offeror’s ability to perform the prospective contract successfully. A well-written response has value only when it accurately represents resources, decisions, dependencies, and risks.
The paradox appears when optimization weakens that connection. The closer AI brings every response to an ideal rhetorical form, the less evaluators can infer from form itself. Strong organization no longer proves that a team has organized the work. Detailed risk language no longer proves that anyone has identified the project’s actual risks.
Research presented through the Naval Postgraduate School’s acquisition symposium describes a related scenario. Industry may use AI to draft proposals while government systems use AI to evaluate them. The resulting human agency problem concerns accountability, fairness, ethical judgment, and institutional knowledge.
This is where the commentary’s second label becomes useful. “Garcia’s law” should not be confused with a statute, regulation, or controlling procurement doctrine. It functions as an editorial principle within the reported argument. As the cost of generating plausible proposal content approaches zero, the cost of verifying that content rises.
That verification burden does not disappear because a response is grammatically clean. It moves to evaluators, contracting officers, technical teams, security reviewers, and eventually program managers. They must test whether cited experience is relevant, proposed personnel are available, assumptions are realistic, and commitments are traceable.
The story therefore marks a change in evidentiary value. Writing remains necessary, but polished writing is becoming a weaker proxy for operational competence. Agencies that continue scoring it as a strong proxy risk rewarding the bidder with the best generation workflow.
Faster Drafting Creates a Slower Verification Problem
AI saves time before submission, but unverified output can transfer more work and risk to the government after submission.
The strongest case for AI in acquisition is speed. Federal procurement includes market research, requirements development, solicitation drafting, question management, proposal analysis, documentation, and contract administration. Many of these activities involve finding and reconciling text across large document sets.
Government Accountability Office research confirms that agencies are already acquiring AI through different contract approaches. Its April 2026 review found that reported federal AI use more than doubled from 2023 to 2024. The acquisition review also identified difficulty obtaining technical experts to evaluate proposals and understanding AI-related costs. Separately, the Office of Management and Budget’s M-25-22 memorandum on federal AI acquisition directs agencies to procure effective and trustworthy AI while addressing competition, data portability, vendor lock-in, and performance monitoring.
Those findings expose the capacity mismatch behind the paradox. AI lets bidders create more complete and sophisticated submissions. However, an agency may not have enough data scientists, acquisition specialists, or mission experts to validate every technical statement efficiently.
Longer answers can intensify the problem. A model can produce extensive implementation detail for every requirement, even where the bidder has little direct experience. Evaluators then face hundreds of claims that sound specific but may rest on generic patterns, outdated records, or unsupported assumptions.
Hallucination is only one risk. A proposal does not need a spectacular falsehood to mislead an evaluator. Smaller distortions can accumulate through embellished past performance, incompatible staffing commitments, optimistic transition plans, or security controls copied from another environment.
Retrieval-augmented generation reduces some errors by supplying approved source material to a model. It does not eliminate the need for review. A retrieved document can be obsolete, irrelevant to the solicitation, or stripped of qualifications that change its meaning.
The danger grows when AI output becomes recursive. One model drafts a technical approach from an internal library. Another scores that approach against the solicitation. A third writes the evaluator’s findings. If all three systems share similar language patterns or model families, apparent agreement may reflect correlated behavior rather than independent confirmation.
Automation bias adds another layer. People tend to defer to a structured output when it arrives with confidence, detail, and numerical scoring. A busy reviewer may examine the model’s conclusion instead of reconstructing the evidence behind it.
The American Bar Association has explored how these risks can reach procurement integrity. Its analysis of automated procurement risk discusses data poisoning, evasion, prompt injection, and automation bias. One scenario involves hidden text directing an evaluation system to rate a proposal favorably.
That example shows why conventional document review is insufficient. A proposal may contain content invisible to a person but readable by a machine. Conversely, an automated extractor may omit charts, qualifications, or formatting that a human evaluator would consider important.
Agencies therefore need controls at both boundaries. Incoming documents require sanitization and inspection before automated processing. Model outputs require citations, logs, reproducible evaluation criteria, and named human owners.
The government also needs to distinguish administrative assistance from substantive judgment. AI can help locate a requirement or identify inconsistent terminology. Deciding whether an approach creates unacceptable mission risk is a different function. That decision should remain attributable to an authorized official.
None of this erases the time savings. It changes where leaders should measure them. Hours removed from drafting do not represent a net benefit if they produce larger review queues, weaker records, or performance disputes later.
AI-Assisted Proposals Put Evaluators Under Pressure
Federal evaluators must replace stylistic signals with evidence that is difficult to generate and easy to verify.
The immediate pressure falls on acquisition teams. They have limited evaluation periods, uneven access to technical specialists, and a duty to treat competitors fairly. They also need a clear record explaining why the government selected one offeror over another.
That record matters beyond the initial award. It supports debriefings, internal reviews, audits, and bid protest litigation. If an agency uses an opaque model to reach a material conclusion, reconstructing the decision becomes difficult.
Federal procurement already provides a human accountability structure. Evaluation boards assess proposals, contracting officers manage the acquisition, and source selection authorities make documented judgments. AI does not automatically fit those roles because a model cannot hold delegated authority or explain its reasoning like a responsible official.
A useful dividing line is whether a tool changes the substance of an evaluation. Search, indexing, formatting, and requirement tracing can support reviewers. Assigning strengths, weaknesses, risks, or comparative value affects the award decision and therefore demands tighter control.
The same distinction applies to contractors. Using AI to format an approved staffing plan differs from asking it to invent one. Summarizing verified past performance differs from generating an idealized project history. Proposal leaders need approval gates that reflect those differences.
Agencies can also redesign what they ask bidders to submit. Long narrative sections are highly compatible with text generation. Work samples, technical challenges, structured data, oral presentations, and scenario responses can reveal more about the people expected to perform.
Live demonstrations are not immune to coaching or automation. They do, however, let evaluators test assumptions and ask follow-up questions. A team that understands its design should explain tradeoffs, failure modes, dependencies, and alternatives without relying on a prewritten answer.
Past performance also becomes more valuable when verified carefully. Instead of accepting broad claims, evaluators can examine whether the cited work involved a comparable mission, operating environment, security boundary, and level of complexity. References should confirm the contractor’s actual role.
Staffing evidence needs similar scrutiny. AI can create an impressive management narrative around resumes that do not support it. Agencies can test availability, labor-category alignment, key-person commitments, and the relationship between proposed personnel and the technical approach.
Cost realism remains essential. A beautifully generated plan may promise aggressive schedules, extensive testing, and continuous oversight without including the labor required to deliver them. Comparing narrative commitments with staffing and pricing can reveal contradictions.
The pressure also extends to smaller agencies. Large departments may build secure evaluation environments and assemble multidisciplinary review teams. Smaller buyers may depend on commercial tools without equivalent legal, technical, and security support.
Shared acquisition services can help, but standardization introduces another concern. If many agencies use the same scoring model, contractors will learn its preferences. Proposal optimization can then target the evaluator rather than the mission.
This is the acquisition version of search engine optimization. Google News ranks and organizes information for discovery, while publishers adapt to its signals. An automated evaluator creates a higher-stakes feedback loop. Contractors will adapt wording, structure, and evidence to whatever the system rewards.
The remedy is not secrecy alone. A hidden scoring process can undermine fairness and make errors harder to challenge. Agencies need transparent criteria while protecting system instructions and monitoring for manipulation.
Good evaluation design makes gaming less valuable. It rewards corroborated facts, measurable commitments, relevant experience, and coherent tradeoffs. It also uses multiple forms of evidence rather than allowing a single narrative score to determine the outcome.
The Real Contest Is Persuasion Versus Proof
AI can improve persuasion on both sides of a competition, but it cannot replace evidence that a contractor can execute.
This is the article’s central reversal. Generative AI appears to make proposal quality more important because every submission can become clearer and more responsive. In practice, it makes conventional proposal polish less decisive.
The old signals do not vanish at once. A confusing proposal can still indicate poor coordination, and an incomplete response can still fail solicitation requirements. The threshold for acceptable presentation simply rises while differentiation moves elsewhere.
Proof begins with traceability. Every significant claim should connect to an approved source, named owner, measurable commitment, or documented assumption. Contractors need to know which statements came from a model and who validated them before submission.
A practical internal record can associate each claim with the solicitation requirement, source document, subject-matter reviewer, and final approval. This resembles a searchable knowledge base, but its purpose is contractual discipline rather than convenience.
Traceability also helps after award. Program teams can see which proposal commitments became contractual obligations and which assumptions require confirmation. Without that continuity, the proposal remains a persuasive artifact disconnected from delivery.
Agencies should ask for evidence in reusable formats. Structured staffing data, milestone definitions, test criteria, security responsibilities, and dependency lists are easier to compare than unrestricted prose. Narrative remains useful for explaining judgment, but it should not obscure the underlying commitments.
Technical demonstrations provide another proof layer. A bidder can show how its proposed architecture responds to a failure, protects government data, or supports portability. Evaluators can then compare observed behavior with written claims.
For AI systems, testing should include imperfect data, adversarial inputs, drift, and operational constraints. A model that performs well during a curated demonstration may fail when users introduce ambiguous requests or sensitive information.
GSA’s proposed rule for large language model systems illustrates the growing focus on operational safeguards. The June 2026 draft LLM clause addresses government data, intellectual property, privacy, portability, change notification, evaluation, and remediation.
The draft also shows why rigid controls can create their own problems. Requirements that diverge sharply from commercial practices can discourage vendors or become difficult to flow down through model providers and subcontractors. Agencies must protect government interests without specifying obligations that no credible supplier can accept.
That is a tradeoff, but it should remain supporting context. The primary conflict is still persuasion versus proof. Disclosure and data safeguards matter because they determine whether agencies can trust and verify the systems involved.
Independent review is another proof mechanism. A person who did not draft the response should test major claims against source records. High-risk sections, including security, staffing, cost, transition, and past performance, deserve stronger review than low-risk formatting.
AI can support this challenge function, but it should not grade its own work. A separate retrieval corpus, model, prompt, or review team can reduce correlated errors. Human specialists must resolve material conflicts.
Contractors also need clear rules about confidential and controlled information. Uploading solicitation data, pricing, resumes, or technical designs to an unapproved service can create privacy, security, and intellectual property exposure.
Agency evaluators face an equivalent duty. Proposal information is sensitive, and placing it into a public model can compromise procurement integrity. Approved environments should define retention, training use, access control, logging, and incident response.
The winning organization will not necessarily be the one that avoids AI. It will be the one that uses automation while preserving a trustworthy chain from source evidence to final commitment.
What Garcia’s Law Cannot Tell Us Yet
The warning is credible, but neither the perfect proposal paradox nor Garcia’s law has been established as a formal procurement rule.
The terminology requires caution. The underlying Google News item points to Federal News Network commentary, not a regulation, court decision, or government-wide policy. Readers should treat its named concepts as analytical shorthand.
There is no public evidence that every AI-assisted proposal is unreliable. Many contractors already use controlled workflows that limit models to approved content and require expert review. AI may improve accuracy by finding omissions, inconsistent figures, or unsupported statements.
There is also no basis for assuming that human-written proposals are inherently trustworthy. Traditional proposal teams can exaggerate capabilities, reuse irrelevant content, or make promises that delivery teams cannot keep. The integrity problem predates generative AI.
AI changes scale, speed, and detectability. It lets teams create more plausible material with less effort. It can also make the origin of a statement difficult to reconstruct unless the organization maintains logs and citations.
Disclosure mandates are not an automatic solution. A simple checkbox asking whether AI was used reveals little because the term covers spell-checking, retrieval, drafting, scoring, and autonomous generation. Overbroad disclosure can also punish benign use without exposing risky practices.
A better approach focuses on materiality. Agencies should care whether AI influenced claims about cost, performance, security, staffing, schedules, or technical capability. They should also ask what validation and approval occurred.
Model detection is unlikely to provide dependable enforcement. Text detectors produce false positives and can be defeated through editing. Evaluators should examine evidence and consistency rather than trying to guess who typed each sentence.
Bias remains another unresolved concern. Historical procurement data may reflect past preferences, unequal access, or inconsistent scoring. Training an evaluation system on that record can reproduce those patterns while presenting them as objective.
Transparency has limits as well. Publishing every model instruction might help contractors manipulate the system. Keeping the process entirely secret could deny bidders a meaningful understanding of how the agency judged them. Agencies need enough disclosure to support fairness without exposing exploitable details.
Data quality may be the hardest constraint. An evaluator cannot reliably compare performance if past records are incomplete, inconsistent, or disconnected across systems. Better models do not repair weak acquisition data by themselves.
The GAO’s call for agencies to collect and apply lessons learned is relevant here. Federal buyers need records of which AI acquisition methods worked, what failed during performance, and how evaluation findings predicted actual outcomes.
This creates a measurable test for the paradox. If increasingly polished proposals correlate less strongly with contract performance, agencies should reduce the weight given to narrative presentation. If structured evidence and demonstrations predict outcomes better, evaluation plans should favor them.
Until that evidence accumulates, agencies should avoid two overclaims. They should not present AI evaluation as neutral simply because it uses software. They should not assume every human judgment is superior simply because a person made it.
The defensible position is narrower. Material decisions require accountable officials, reviewable evidence, and a process capable of detecting error or manipulation. AI can assist that process, but it does not satisfy those conditions automatically.
Three Signals That Will Test the Paradox
The next phase will be determined by procurement design, protest records, and evidence from contract performance.
The first signal is a change in solicitation structure. Watch whether agencies reduce long narrative requirements and add technical challenges, oral presentations, structured commitments, or demonstrations. That shift would strengthen the argument that conventional prose has lost value as a differentiator.
The details will matter. Replacing a written essay with a scripted presentation changes little. A meaningful test asks proposed personnel to explain decisions, respond to new facts, or demonstrate a capability under realistic constraints.
The second signal is the first visible dispute over material AI use in proposal evaluation. A protest, audit, or inspector general review could reveal whether an agency relied on automated findings and whether its human officials exercised independent judgment.
The central questions will concern the record. Did the tool generate a rating or merely organize evidence? Could evaluators reproduce its output? Did they examine contrary information? Who accepted responsibility for the final judgment?
A well-documented case would strengthen confidence in bounded AI assistance. An opaque process that cannot explain a material rating would reinforce the paradox and increase pressure for government-wide controls.
The third signal is performance data. Agencies should compare proposal claims with staffing stability, delivery milestones, cost growth, security incidents, and user outcomes after award. That evidence can show which evaluation signals actually predict success.
Performance measurement will also test contractors’ internal controls. Organizations that connect generated content to verified records should experience fewer contradictions during delivery. Those that optimize only for submission may struggle when program teams inherit unrealistic commitments.
Google News will continue surfacing debates about AI policy, federal contracts, and automated work. Readers should look past the aggregation label and ask whether agencies are changing what they reward.
The best response is neither an AI ban nor unchecked automation. Federal buyers should make polished language the entry point, not the finish line. They should demand evidence, test key claims, protect sensitive data, and document accountable human decisions.
For contractors, the same principle applies before submission. Build a reviewable chain from source records to proposal claims, then preserve it for delivery. Teams managing large collections of evidence can use knowledge blending to organize context, but accountable experts must still approve every material promise.
What would prove this approach is working? Watch for shorter narratives, harder demonstrations, clearer audit trails, and better links between evaluation findings and contract results. If those signals appear, the perfect proposal paradox will have forced a useful correction. If they do not, AI may make federal proposals easier to produce while leaving successful acquisition just as difficult to achieve.


