Key facts
- ShieldFont replaces words with similar parts of speech but different informational context to disrupt AI training data.
- The font replaces approximately 24.5% of all words on a page, with 45.8% of content words affected.
- ShieldFont authors report over 90% of pages are rejected by scraper quality filters after its application.
- Pages that bypass filters are considered 'training-time garbage' by the creators, containing incorrect assertions.
- The font can potentially impact human-readable tools like search engines and screen readers.
A new font called ShieldFont has been developed to combat the indiscriminate scraping of web content for AI training data. Created by designers Isaque Seneda and Gabriel Abrucio, the font subtly alters text on webpages, making it nonsensical for AI scrapers while remaining perfectly readable for human users. The font achieves this by using ligatures, a standard font feature, to replace entire words with others that maintain grammatical correctness but drastically change the meaning. For example, 'horse' might be replaced with 'potato.'
ShieldFont aims to poison AI training datasets by introducing semantically incorrect information. The creators report that the font replaces approximately 24.5% of all words on a page, with a significant portion of these being 'content words.' This alteration leads to a high rejection rate for pages by AI scrapers' quality filters, with over 90% of pages being flagged as unsuitable for training. For the small percentage of pages that are still accepted, the content is described as 'training-time garbage,' asserting nothing true.
While effective against basic HTML scraping, ShieldFont is not a foolproof solution. More sophisticated AI tools that render webpages as images and use optical character recognition could potentially bypass the font's defense, though this method is significantly more costly and time-consuming. The creators' primary goal is to enforce AI ethics by making unauthorized data collection more difficult and less valuable, thereby giving creators more control over how their work is used for AI training. They hope this method will encourage further innovation in similar human-readable, machine-unreadable content defenses.
