top of page

Decoy Font Hides Text From AI, but Its Spatial-Frequency Trick Has an Expiration Date

Jul 17
13 min read

Updated: Jul 20

Decoy Font now lets anyone place two messages inside one installable typeface, despite AI systems expecting each character to have one readable identity. The foreground shows a crisp decoy letter. A blurred shape behind it carries the intended letter at a lower spatial frequency. Mixfont says current multimodal models often report the decoy and miss the message that a person can reveal by stepping back or squinting.

That conflict makes Decoy Font more than an unusual typography project. It converts a familiar optical illusion into an adversarial interface between people and machine vision. The font asks which visual scale an AI system trusts, then supplies a plausible but incorrect answer at that scale.

The experiment also exposes a hard limit. Decoy Font hides text from AI only when the model follows a vulnerable reading process. Resizing, blurring, image preprocessing, informed prompting, or access to the underlying text can reveal the hidden layer. The result is a useful demonstration and a possible scraping deterrent, but it is not encryption.

Decoy Font Turns Every Character Into a Visual Dispute

The important change is not that a strange font can confuse OCR, but that a standard font file can encode competing readings inside every glyph.

Mixfont describes Decoy Font as a downloadable TrueType font, or TTF, that prints a decoy for each letter. Users can install it, type ordinary text, and copy the underlying characters into another application. They can also use Mixfont’s browser playground to pair a hidden message with a decoy message and export the result as an image.

A conventional typeface maps each stored character to one visible glyph. Decoy Font preserves that basic software relationship while making the rendered glyph perceptually ambiguous. One component contains thin, sharply defined outlines. The other uses a broader, blurred mass that carries a different letter.

A viewer examining the image closely tends to notice the narrow outlines first. A viewer who moves farther from the display loses some fine detail, allowing the broad background shape to dominate. Squinting produces a related effect because it suppresses high-frequency detail.

That creates two readings without animation, interactive code, or a custom image-processing application. The same rendered area can communicate one sequence nearby and another sequence from a distance. The hidden phrase remains tied to the actual typed characters, while the visible decoys occupy the sharper layer.

Mixfont says it tested generated images with advanced multimodal systems and observed models reading the foreground decoy. Its page includes examples in which systems reportedly return “BINGE SHOWS” instead of the intended message. Those screenshots document individual trials, not a controlled benchmark across models, settings, and preprocessing methods.

The distinction matters. A demonstration proves that a failure mode exists under specific conditions. It does not establish a stable boundary between human and machine perception. Image resolution, compression, font size, contrast, prompt wording, and model updates can all change the outcome.

The project follows Mixfont’s earlier Ghost Font experiment, which hides letters in the coherent motion of dots. Ghost Font requires video because no single frame cleanly contains the intended message. Decoy Font removes that dependency and packages its illusion inside a static, installable font.

That shift makes the new experiment easier to distribute. A designer can apply a TTF in a familiar graphics application. A creator can export a social image without building an animation. A developer can test the result using the same screenshot workflow that feeds text into a vision model.

It also broadens the pressure test. Ghost Font asks whether a model can integrate motion across frames. Decoy Font asks whether a model can inspect one image at more than one visual scale. Both exploit a mismatch between the information available in the media and the information selected by the recognition pipeline.

The experiment arrived as multimodal assistants became common tools for reading screenshots, documents, interfaces, and social posts. Their convenience depends on turning pixels into text quickly. Decoy Font challenges the assumption that the most obvious extracted string is the only string present.

How Decoy Font Uses Spatial Frequency to Split One Image

Decoy Font works because a letter’s broad silhouette and its fine edges occupy different bands of visual information.

Spatial frequency describes how quickly brightness or color changes across an image. High spatial frequencies represent narrow edges, small details, and rapid transitions. Low spatial frequencies represent broad shapes, gradual shading, and large areas of contrast.

Decoy Font assigns different letters to those two bands. The decoy occupies the high-frequency layer through thin outlines. The hidden letter occupies the low-frequency layer through a blurred, heavier form. Both appear in the same location, but viewing conditions determine which layer controls perception.

The mechanism comes from research on hybrid images. In a widely cited 2006 paper, Aude Oliva, Antonio Torralba, and Philippe Schyns described images whose interpretation changes with viewing distance. Their hybrid image research combines the low-frequency content of one picture with the high-frequency content of another.

The familiar Einstein and Marilyn Monroe illusion demonstrates the principle. Einstein’s fine facial details dominate at close range. At a greater distance, those details become harder to resolve, and Monroe’s broad facial structure becomes more apparent.

Decoy Font replaces faces with glyphs. That substitution looks simple, but letters impose stricter constraints. A useful glyph must remain recognizable beside other glyphs, preserve spacing, and survive common rendering systems. Each decoy also needs enough internal room to contain a distinct low-frequency shape.

The result depends on scale rather than secrecy. No hidden key controls the transformation. The intended letter already exists in the pixels. A viewer reveals it by changing the effective resolution of the visual signal.

Downsampling can reproduce the effect computationally. When an image becomes smaller, thin contours lose definition while broad shapes survive. A blur filter can also remove the high-frequency outlines and leave the lower-frequency mass. Even viewing a thumbnail can alter which message appears dominant.

This explains why the font can fool one AI request yet fail another. A vision pipeline might resize an uploaded image before recognition. Another might crop the text and enlarge it. A third might run several OCR passes at different resolutions. Each decision changes the balance between the two messages.

Traditional optical character recognition, or OCR, converts images of text into machine-readable character sequences. Modern multimodal models often combine visual encoders with language models, allowing them to interpret layouts and answer questions about screenshots. That added reasoning does not guarantee careful multiscale inspection.

A model can produce a fluent answer after extracting the wrong layer. Language modeling may even make the error more persuasive because the system prefers a coherent phrase over an uncertain collection of shapes. A well-chosen decoy exploits both visual selection and linguistic confidence.

The technique differs from hiding white text on a white background or placing instructions in metadata. Both Decoy Font messages are visible features of the rendered image. The ambiguity comes from perception, not from information stored outside the image.

It also differs from cryptography. Encryption transforms information so that an unauthorized observer cannot recover it without a key. Decoy Font leaves the intended letter visually recoverable and offers the observer a clue: every glyph has an unusual dual structure.

That distinction defines the experiment’s real value. Decoy Font demonstrates a machine-vision blind spot using a format people already understand. Its accessibility as a TTF makes the concept easy to test, copy, and eventually defeat.

Why AI Readers Can Accept the Wrong Message

Decoy Font pressures AI systems at the point where visual recognition becomes a confident language answer.

A multimodal assistant usually does more than run a single classic OCR engine. It may resize an image, divide it into patches, encode visual features, locate text, and use linguistic context to infer the final sequence. The exact pipeline varies across products and remains partly undisclosed.

Decoy Font does not need to defeat every stage. It only needs the system to prioritize the outlined decoy before another stage examines the broad hidden form. Once the model produces a plausible phrase, later reasoning can reinforce that initial interpretation.

The attack therefore resembles targeted adversarial text more than ordinary visual noise. Researchers Congzheng Song and Vitaly Shmatikov showed in 2018 that carefully modified text images could make OCR produce attacker-selected words. Their OCR attack study generated examples that changed recognized words while preserving a different reading for people.

That work also established why targeted errors matter. Random gibberish announces that recognition failed. A fluent but incorrect phrase can pass through downstream classifiers, databases, search systems, or automated decisions without triggering the same suspicion.

Decoy Font supplies a full competing character rather than a nearly invisible mathematical perturbation. Its deception is apparent once someone understands the construction. However, the false reading can remain semantically clean, especially when users choose decoy words that form a natural sentence.

Consider a creator posting a graphic that says “DRAFT RESULTS” in the hidden layer and “DINNER PLANS” in the outlined layer. A person who knows to step back can recover the intended phrase. An automated indexer might store the harmless decoy and make the graphic harder to find through the sensitive terms.

The opposite scenario creates a security concern. A document might look acceptable to a distant viewer while an automated workflow extracts a different instruction from its sharp details. A system that trusts screenshot OCR could then summarize, classify, or route the document using the decoy.

The underlying text adds another complication. Because Decoy Font is a real font, the typed character and the visually emphasized character can differ in practice. Copying text from a document may expose the encoded characters directly, bypassing the visual puzzle. A browser scraper reading the document object model would not need OCR at all.

That means the font mainly targets pixel-based reading. It has a stronger chance against screenshots, flattened graphics, image-only posts, and visual document pipelines. It offers little protection when an agent can inspect HTML, extract PDF text, query accessibility labels, or access the source document.

Even within image-only workflows, a capable agent has several options. It can request multiple crops, invert colors, blur the image, examine lower-resolution copies, or ask a coding tool to analyze frequency bands. A prompt that explicitly mentions hybrid images can direct attention toward the hidden layer.

Mixfont acknowledges this limit. The project’s description says agents with coding abilities can see past the initial lettering and that informed prompting can reveal hidden letters. The claim is deterrence against basic recognition, not guaranteed exclusion of AI.

That caution separates a useful experiment from a privacy product. A vision failure observed today can disappear after a silent model update. The same model can also behave differently when an application changes its image preprocessing.

The current evidence lacks a public test corpus containing many font sizes, hidden-decoy pairs, resolutions, and compression levels. It also lacks repeatable measurements across open OCR engines and commercial vision models. Without those controls, nobody can assign a reliable success rate to Decoy Font.

Still, the failure has diagnostic value. It shows that fluent multimodal answers can conceal uncertainty about perceptual scale. A model that reports only one message may be less observant than its confident wording suggests.

Decoy Font Hides Text From AI, Not From a Determined System

The font is best treated as a speed bump for automated reading, because its defensive effect disappears once the observer tests multiple scales.

The simplest countermeasure is image transformation. A system can create several resized versions, apply blur at increasing strengths, and run recognition on each output. If the extracted strings change, the image should be flagged as ambiguous rather than assigned one authoritative transcript.

Frequency filtering offers a more direct response. A low-pass filter suppresses fine contours and preserves broad masses. A high-pass filter emphasizes edges and removes gradual shading. Running OCR on both outputs would expose the two channels that Decoy Font deliberately combines.

This is not a speculative defense requiring a new foundation model. Image libraries already support resizing, Gaussian blur, edge extraction, and Fourier analysis. An agent with permission to execute code can arrange these operations automatically after detecting unusual glyph structure.

The font is also public. Defenders can download samples, generate letter pairs, and include them in evaluation or training data. Adversarial training, which exposes a model to deliberately difficult inputs, can teach recognition systems to treat multiscale disagreement as a warning signal.

Research into OCR defenses shows why no single fix should inspire too much confidence. Song and Shmatikov noted that OCR has a vast output space because it must recognize arbitrary character sequences. They also found that visual similarity does not imply semantic similarity. Tiny changes can transform a word into another word with an opposite meaning.

A 2025 study on visual OCR defenses examined adversarial fonts, color distortions, animated interference, and dynamic blur. That broader category confirms continuing interest in making text legible to people but harder to extract automatically.

Yet resistance to one OCR configuration does not establish resistance to all machine readers. Different models use different input sizes and features. An effect that survives one pipeline can vanish after another pipeline rescales or sharpens the image.

Human readers also vary. Distance, display density, visual acuity, contrast sensitivity, lighting, and font size influence which layer becomes visible. A hidden message that appears obvious to its designer may remain illegible to another person.

That variability creates an accessibility problem. Low-vision users, people with contrast-sensitivity differences, and screen-reader users may not receive the intended message. If the image includes an accurate text alternative, automated systems can read that alternative too. If it omits the alternative, some people lose access.

The World Wide Web Consortium recommends providing equivalent text for images that convey meaningful words. Its image text guidance also favors real text when presentation can be achieved without an image. Decoy Font’s purpose creates direct tension with those practices.

CAPTCHA is therefore a risky application, even though Mixfont mentions it as a possible direction. A visual challenge that depends on subtle frequency perception can exclude legitimate users. The W3C’s CAPTCHA accessibility note warns that common robot tests can also block people with visual, auditory, or cognitive disabilities.

The concern extends beyond accessibility. A decoy font can support defensive obscurity, but it can also manipulate document-processing systems. An invoice, policy notice, or contract that produces conflicting human and machine readings poses an integrity risk.

Organizations should not treat the font as a safe way to distribute confidential material. If disclosure would cause harm, the content needs access control, secure storage, and encryption. A visual illusion does not revoke copies, authenticate recipients, or prevent manual transcription.

For personal research and sensitive notes, the safer pattern is to keep source material inside a controlled personal knowledge system. Visual obfuscation can supplement that boundary, but it should not become the boundary itself.

The security lesson runs in both directions. Creators should not assume an AI failure is permanent. AI product teams should not assume a plausible OCR transcript reflects everything a person can see.

The Real Pressure Falls on Multimodal AI Evaluation

Decoy Font matters most as a benchmark idea because it tests whether an AI system recognizes ambiguity before acting on extracted text.

Most public evaluations ask whether a model returns the expected answer. Decoy Font introduces a harder requirement: the model should notice that the image supports two answers at different scales. Reporting either string without qualification is an incomplete reading.

A stronger evaluation would generate a balanced corpus of hidden and decoy pairs. It would vary character combinations, phrase lengths, point sizes, image dimensions, colors, contrast, compression, and background texture. It would also test screenshots from different operating systems because font rasterization changes the pixels.

Researchers could then measure three outcomes. First, does the system recover the intended low-frequency message? Second, does it recover the high-frequency decoy? Third, does it explicitly report the conflict instead of pretending the image contains one uncontested string?

That third measure is the most revealing. An assistant that says “I see one phrase in the outlines and another after reducing detail” behaves more safely than one that confidently selects either layer. Recognition accuracy alone misses this calibration problem.

The benchmark should include transformations available to ordinary users. Tests could compare the original image with thumbnails, blurred copies, enlarged crops, grayscale versions, and printed scans. Each transformation would reveal how stable the effect remains outside Mixfont’s playground.

Open OCR systems belong in the comparison alongside multimodal assistants. Their results would help separate language reasoning from visual extraction. If a classic OCR engine reads the decoy while a vision-language model finds both layers, the reasoning system adds useful analysis. If both fail identically, preprocessing may be the common weakness.

Human testing is equally important. Participants should not receive instructions to squint or step back during the first trial. Researchers could then test discovery, legibility, reading speed, and error rates across viewing conditions. Accessibility groups should be included rather than treated as an afterthought.

This kind of work would place Decoy Font beside other adversarial evaluations instead of presenting it as a permanent anti-AI shield. Its greatest contribution may be a clean, intuitive illustration of multiscale disagreement.

The project also pressures applications that turn screenshots into actions. An assistant reading a meme has limited consequences. An agent reading a button label, purchase order, or policy image can initiate a workflow based on the wrong text.

Product teams can respond by separating recognition from authorization. Text extracted from an image should remain untrusted input, especially when it triggers payments, data changes, or external messages. Multiple recognition passes and human confirmation can reduce the risk.

Search platforms face a related challenge. If visually ambiguous images become common, indexing one transcription can misrepresent their content. Storing alternative readings or an ambiguity score would be more honest, although it would complicate ranking and moderation.

Content moderators also need to consider both layers. A harmless decoy can conceal prohibited material from automated filters. Conversely, an aggressive decoy could cause a benign hidden message to receive an incorrect moderation label.

These risks do not mean Decoy Font itself is dangerous. They show that visual language models operate inside larger systems that often assume text extraction is objective. The font makes that hidden assumption visible.

The pressure therefore falls on evaluation designers, agent developers, accessibility reviewers, and security teams. Each group needs to decide what happens when people and machines do not agree about the words in an image.

Three Signals Will Show Whether the Trick Lasts

Decoy Font’s future depends on repeatable testing, automated multiscale recovery, and evidence that people can read it without unacceptable friction.

The first signal is a reproducible benchmark. Mixfont or independent researchers would need to publish image sets, prompts, preprocessing settings, model versions, and expected readings. If several current systems continue selecting the decoy across varied conditions, the case for a meaningful blind spot strengthens.

If results collapse after small changes in resolution or prompting, the opposite conclusion follows. The font would remain an engaging optical demonstration, but claims about protection from AI scraping would weaken. Repeatability matters more than a collection of successful screenshots.

The second signal is automatic recovery by mainstream assistants. Watch for models that independently resize, blur, or filter suspicious text without being told about Decoy Font. A system that reports both messages has effectively neutralized the deception, even if it cannot determine which message the author intended.

This response can arrive through model training or through agent tools. A vision model does not need to internalize every optical illusion if its surrounding system can launch a basic image-analysis routine. Tool use shortens the likely lifespan of static visual defenses.

The third signal is human usability data. Decoy Font needs testing across devices, ages, vision conditions, and reading distances. If intended readers require repeated hints, unusually large text, or ideal display conditions, the format will remain unsuitable for routine communication.

Accessibility performance will be especially important if developers explore CAPTCHA, authentication, or public information uses. A method that slows AI by also excluding people has not preserved a useful human advantage. It has merely moved the failure.

These signals point toward a restrained conclusion. Decoy Font hides text from AI under some reported conditions, and it does so through a clever application of established perceptual research. It does not create a durable private channel.

That does not make the experiment trivial. It offers creators a tangible way to probe the limits of screenshot readers. It gives researchers an understandable adversarial test. It reminds AI users that a confident transcript can represent one sampling of an image rather than its complete meaning.

Try the font only with information you can afford to expose. Compare the original image with resized and blurred copies, then test several recognition systems. Most importantly, ask whether the system identifies the ambiguity rather than whether it guesses your preferred message once.

The useful question is no longer whether Decoy Font can fool an AI today. It is whether AI developers will teach their systems to distrust the first readable layer before ambiguous visual text reaches search, moderation, or autonomous action.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page