How does a Drupal page tell machines that its content was made with AI?

Deep dive into the AI Disclosure project
A Drupal page tells machines about AI involvement with two elements in the head: a JSON-LD block carrying the schema.org property digitalSourceType, whose value comes from the IPTC Digital Source Type vocabulary, and an ai-disclosure meta tag with the grade, the verdict and the icon. The AI Disclosure module writes both from the grade assigned to the content.

When an AI label is mandatory is covered in the article on Article 50(4) of the AI Act in Drupal. This one looks at the part a machine reads: the property, the value, where it sits in the page and what writes it. The module is AI Disclosure, which I created on drupal.org from the Drupal AI Initiative plan.

How does a Drupal page declare that its content was made with AI?

With two elements in the <head>: a JSON-LD block carrying the schema.org property digitalSourceType, and an ai-disclosure meta tag. AI Disclosure writes both on the canonical page of every piece of content with a resolved disclosure, starting from the grade the editors gave it.

This is the block the site outputs today, in production, on the article Which queries slow down a Drupal route:

{
  "@context": "https://schema.org",
  "@type": "Article",
  "@id": "https://giorgiopagano.org/en/articles/which-queries-slow-down-a-drupal-route",
  "digitalSourceType": {
    "@id": "https://schema.org/CompositeWithTrainedAlgorithmicMediaDigitalSource"
  },
  "disambiguatingDescription": "This content was produced by AI and edited by a person."
}

And this is the meta tag on the same page, identical in both languages:

<meta name="ai-disclosure" content="ai_assisted_hitl; required=not-required; icon=ai_modified">

Nobody writes either of them by hand. The editor picks a grade in the content's field, and the module derives the rest.

What is digitalSourceType and which values does it accept?

digitalSourceType is a schema.org property that states the nature of the digital source of a CreativeWork. The expected value is a member of the IPTCDigitalSourceEnumeration, which mirrors the IPTC Digital Source Type vocabulary. Schema.org lists seventeen members, and according to the Google figures quoted on the same page the property is used on fewer than a thousand domains.

Only a few of those values matter for text:

schema.org member IPTC definition
TrainedAlgorithmicMediaDigitalSource Created using Generative AI
CompositeWithTrainedAlgorithmicMediaDigitalSource Edited using Generative AI
CompositeSyntheticDigitalSource Composite including generative AI elements
AlgorithmicallyEnhancedDigitalSource Algorithmically-altered media
AlgorithmicMediaDigitalSource Pure algorithmic media, with no training data

The property's domain is CreativeWork, so it applies to Article, ImageObject, VideoObject and every other type below it.

Why do schema.org and IPTC use different names for the same values?

In 2024 IPTC retired minorHumanEdits and digitalArt, replacing them with humanEdits and digitalCreation. Schema.org still lists the two retired terms and has neither of the new ones.

IPTC term Status in IPTC schema.org member
minorHumanEdits retired MinorHumanEditsDigitalSource
humanEdits active missing
digitalArt retired DigitalArtDigitalSource
digitalCreation active missing

Whoever writes the markup has to pick between an outdated schema.org name and an IPTC term that schema.org lacks. AI Disclosure handles it like this: the grade stores the IPTC URI, and the module turns it into the schema.org member when one exists. When none does, it outputs the IPTC URI as it is, in the form http://cv.iptc.org/newscodes/digitalsourcetype/humanEdits. The http:// scheme is the one IPTC declares, and with https:// the identifier would name a different resource.

The nine grades shipped with the module only use terms schema.org knows. The IPTC URI branch is there for grades a site creates on its own.

Does the IPTC vocabulary apply to text?

IPTC defines its scheme with the sentence "Indicates from which source a digital image was created": it was designed for images. Schema.org attaches it to any CreativeWork, text included, and Google documents it for forum posts, which are text. Using it on an article is therefore allowed, with definitions written with photos and video in mind.

The trickiest case is the ai_assisted_hitl grade, the one this site uses. The module assigns it compositeWithTrainedAlgorithmicMedia, which IPTC calls "Edited using Generative AI" and explains through inpainting and outpainting. The module itself describes the grade in two ways. The documentation says "A person led, AI helped throughout", and on that reading the composite term fits. The grade's configuration says "AI wrote it, a person directed it", and the sentence shown to readers is "This content was produced by AI and edited by a person": on that reading the closer match is trainedAlgorithmicMedia, "Created using Generative AI".

The URI is a field on the grade, editable at /admin/config/content/ai-disclosure/grades. The issue that defined the marking says so explicitly: vocabulary corrections are configuration updates, with no code change.

For machine translation the ai_translated grade uses algorithmicallyEnhanced, which IPTC defines as a change made by an algorithm "without changing the main content of the media".

Which AI Disclosure grade becomes which value, and with which type?

Each grade is a configuration entity with an IPTC URI. This is the mapping of the nine grades shipped with release 1.0.0-alpha2:

Grade Grade's IPTC URI digitalSourceType output
human_only none property left out
ai_metadata none property left out
ai_rated none property left out
ai_translated algorithmicallyEnhanced AlgorithmicallyEnhancedDigitalSource
ai_summarized compositeWithTrainedAlgorithmicMedia CompositeWithTrainedAlgorithmicMediaDigitalSource
ai_partly_assisted compositeWithTrainedAlgorithmicMedia CompositeWithTrainedAlgorithmicMediaDigitalSource
ai_assisted_hitl compositeWithTrainedAlgorithmicMedia CompositeWithTrainedAlgorithmicMediaDigitalSource
ai_generated_autonomous trainedAlgorithmicMedia TrainedAlgorithmicMediaDigitalSource
ai_deepfake trainedAlgorithmicMedia TrainedAlgorithmicMediaDigitalSource

The three grades without a URI cover content a person wrote, where AI only picked the tags or gave a score. For that content the block carries disambiguatingDescription alone, with the grade's sentence.

The @type depends on the entity. A node becomes Article. A media item becomes ImageObject, VideoObject or AudioObject depending on its source plugin, and MediaObject for any other plugin. Every other entity becomes CreativeWork. The module reads the plugin by its ID, with no dependency on the media module.

What does the ai-disclosure meta tag carry?

The meta tag carries three values separated by semicolons: the grade's machine name, the verdict on whether a label is required, and the icon the reader sees on the page.

<meta name="ai-disclosure" content="ai_summarized; required=undetermined; icon=ai_generated">

The verdict is required, not-required or undetermined, spelled out. A zero would mix up "assessed, out of scope" with "never assessed", and that distinction is the reason the signal exists. When a label is due and an evidently artistic work limits how it is shown, the value gains modality=unobtrusive. The icon is none when the grade names one the module cannot draw: the tag states what the reader sees.

Name and structure come from the Drupal AI Initiative plan, which in issue #3586647 (opens in a new tab) calls the meta tag "a cheap stable secondary signal". The plan had required=0 and required=1: the spelled-out verdict came later, and the module documentation flags it as a change to a published format.

There is also a proposal by David E. Weekly, discussed in the W3C AI Content Disclosure Community Group, that uses the same ai-disclosure name with different content: one of the values none, ai-assisted, ai-generated, autonomous or mixed. According to the explainer, Chromium has a feature entry and Mozilla and WebKit have an open request for a position. A program that reads this meta tag can tell the AI Disclosure format by the semicolon and the required= key.

Why does each translation have its own disclosure?

AI Disclosure keeps one disclosure per entity and per language, in the ai_disclosure field. A translation made with AI is a step the original never went through, and the translated page declares it on its own.

This article shows it on itself. The Italian version inherits the site's editorial profile, with the ai_assisted_hitl grade. The English version uses the reviewed translation profile, and its <head> carries a different value:

"digitalSourceType":{"@id":"https://schema.org/AlgorithmicallyEnhancedDigitalSource"}
<meta name="ai-disclosure" content="ai_translated; required=not-required; icon=ai_modified">

The ai_translated grade becomes AlgorithmicallyEnhancedDigitalSource, and the verdict stays not-required because a person reviewed the translation.

How do you get from an IPTC URI back to the grades?

With the ai_disclosure.grade_lookup service, which runs the marking in reverse: it takes an IPTC URI and returns the grades that carry it. It is meant for a module that finds the URI inside a file, for example in the C2PA manifest of an uploaded image, and has to suggest a grade.

$lookup = \Drupal::service('ai_disclosure.grade_lookup');
$grades = $lookup->gradesForIptcUri(
  'http://cv.iptc.org/newscodes/digitalsourcetype/trainedAlgorithmicMedia',
  \Drupal\ai_disclosure\GradeLookupFormat::Data,
);

On this site, with the grades shipped by the module, the result is:

{
    "ai_generated_autonomous": {
        "id": "ai_generated_autonomous",
        "severity": 70,
        "label": "Autonomously AI generated",
        "label_required": true
    },
    "ai_deepfake": {
        "id": "ai_deepfake",
        "severity": 80,
        "label": "AI deep fake (media)",
        "label_required": true
    }
}

One URI can match several grades, and the service returns all of them, ordered by severity. Picking the grade is up to the calling module, which knows the context of the file. The URI comparison is exact: different spacing or case returns an empty result. The result depends on the grade configuration, and getCacheTags() returns the tags to attach to whatever is built from it.

Does Google read digitalSourceType on an article?

Google Search Central documentation recommends digitalSourceType on DiscussionForumPosting and Comment. For those types Google supports two values: TrainedAlgorithmicMediaDigitalSource for content from a trained model such as an LLM, and AlgorithmicMediaDigitalSource for content from a simpler process, such as a reply bot. For those types, Google treats content without the property as written by a person.

Google's page on Article structured data mentions digitalSourceType zero times. The property is valid for schema.org, and the Google documentation I read only covers it on forums and comments. My view is that the block is worth adding anyway, because it has three audiences: Google, aggregators and the agents that read the page. Any effect on rankings is a hypothesis still to be tested, and whoever promises one today is going beyond the sources.

How do you turn the marking on and check it?

The JSON-LD block and the meta tag are on by default after installation. To check them:

  1. Open /admin/config/content/ai-disclosure/output and make sure both checkboxes, one for the

JSON-LD block and one for the meta tag, are ticked.

  1. Give a piece of content a grade, through a profile or in the AI Disclosure field on the edit

form.

  1. Open the content's canonical page and look for the two elements in the source. On the

English version of the slow queries article the command is:

curl -s https://giorgiopagano.org/en/articles/which-queries-slow-down-a-drupal-route | grep -o '"digitalSourceType":{[^}]*}\|<meta name="ai-disclosure"[^>]*>'

The expected result is two lines, the digitalSourceType value and the meta tag:

"digitalSourceType":{"@id":"https://schema.org/CompositeWithTrainedAlgorithmicMediaDigitalSource"}
<meta name="ai-disclosure" content="ai_assisted_hitl; required=not-required; icon=ai_modified">

For another site, change the URL. With a grade that has no IPTC URI the first line is missing, which is the correct behaviour.

  1. Paste the source into the schema.org validator (opens in a new tab) and check that

the Article has digitalSourceType with no errors.

Media only get the marking when they have a page of their own. Drupal ships media.settings.standalone_url turned off, and without that setting /media/1 has no page.

How does it coexist with the theme's JSON-LD?

On this site an article page carries two Article blocks: the theme's, with title, author and date, and the module's, with digitalSourceType. Both use the same @id, the canonical URL, and in JSON-LD the same @id means the same resource.

This is how the two validators handle them. I ran the public URL of the slow queries article through both:

Tool Article items Result
Schema Markup Validator 1, with the properties of both blocks 0 errors, 0 warnings
Google Rich Results Test 1 valid, with digitalSourceType next to headline 2 optional issues on datePublished, which comes from the theme

The merge works because the theme uses the same canonical URL as its @id as the module does. A theme with a different @id produces two separate objects, and digitalSourceType ends up on an Article with no title.

What does the marking leave out?

The marking is a page-level signal, declared by the publisher. It lives in the HTML: an image saved from the page loses it. Marking the file itself, with IPTC in the XMP metadata or with a signed C2PA manifest, is out of scope for version 1.0 of the module, and Article 50(2) assigns it to the provider of the AI system. The page marking records what the editors declared, and it is as reliable as they are.

Sources

All sources were checked on 27 and 28 September 2026.

  1. digitalSourceType (opens in a new tab), schema.org. Domain CreativeWork and use on fewer than a thousand domains according to Google figures. Primary source.
  2. IPTCDigitalSourceEnumeration (opens in a new tab), schema.org. The seventeen members of the enumeration. Primary source.
  3. Digital Source Type (opens in a new tab), IPTC NewsCodes. Term definitions, creation and retirement dates. Primary source.
  4. Discussion forum structured data (opens in a new tab), Google Search Central. digitalSourceType on DiscussionForumPosting and Comment, the two supported values, the default when it is missing. Primary source.
  5. Machine-readable marking: JSON-LD and meta output with IPTC digitalSourceType (opens in a new tab), issue #3586647 of the Drupal AI Initiative. The meta tag format and the rule that leaves vocabulary corrections to configuration. Primary source.
  6. AI Content Disclosure for HTML (opens in a new tab), explainer by David E. Weekly for the W3C AI Content Disclosure Community Group. A proposal, with the implementation status reported by the explainer.
  7. AI Disclosure source code (opens in a new tab), git.drupalcode.org, release 1.0.0-alpha2. src/MachineReadableMarking.php and the grades in config/install: first-hand code.
  8. Schema Markup Validator (opens in a new tab), schema.org, and Rich Results Test (opens in a new tab), Google. Test on the public URL of the slow queries article.
  9. Article structured data (opens in a new tab), Google Search Central. No occurrence of digitalSourceType. Primary source.
  10. AI Disclosure project page (opens in a new tab), drupal.org. Created by sjpagan, release 1.0.0-alpha2.
  11. Grades documentation (opens in a new tab), AI Disclosure. The "Pick it when" column for each grade. Primary source.
Giorgio Alfredo Pagano
AI modified

This text was translated by AI.

How was AI used?

Translated from the Italian original with AI assistance, then read and corrected by a person, who holds editorial responsibility.