An Open Inquiry to Substack Inc.
Questions about the data practices and publisher controls behind Scan for AI text
To Substack Inc.:
I write as a writer and publisher on Substack following Substack’s introduction of Scan for AI text for posts and Notes published on or after July 21, 2026, as well as comments and replies.
I am writing to request an integration-specific explanation of the data flow, contractual terms, publisher controls, retention practices, dispute procedures, and support pathways governing this feature.
The public documentation states several protections. It does not yet provide a complete account of what happens when a publisher’s writing is submitted for analysis.
I am asking Substack to clearly document the relevant processing so that I can understand the system and make informed decisions about whether and how to publish my work on Substack.
Some of the questions below concern processing performed by Pangram. I am addressing this inquiry to Substack alone because Substack selected the provider, is party to the agreement governing the processing, and is the company with which I have a publishing relationship. The public materials do not explain whether or how Pangram’s terms apply to me when my work is processed through the integration. That issue is itself among my questions. Where an answer requires information or confirmation from Pangram, please obtain and include it in Substack’s response.
What I found in the public documentation
Substack’s help-center article states that Scan for AI text uses Pangram to estimate whether eligible text is human-written or AI-assisted. It states that neither company uses publisher content to train generative-AI models.
The article explains that, to disable detection on a post or Note, a publisher must first generate a Pangram analysis and then select Disable detection from the resulting report. Readers will thereafter see an “AI detection unavailable” message. The article does not explain why generating an analysis is a prerequisite to disabling detection, for whose benefit that initial analysis is performed, who may access or use the result, or whether disabling detection prevents future transmission to Pangram rather than merely suppressing the reader-facing result.
The same article instructs users to submit feedback through Report detection error and states that the feedback will be reviewed to “improve detection quality.”
Pangram’s Privacy Policy separately states that submitted data is not used to train, develop, refine, or improve AI models or any machine-learning or AI systems, and is not used to build new products or expand the capabilities of existing products. The policy does permit Pangram to access submission data to diagnose and fix reported issues and to process data for troubleshooting, debugging, security, and maintenance.
The public documentation does not explain how detection-error feedback is reviewed, whether the underlying content or analysis accompanies the feedback, or how “improve detection quality” is distinguished from developing, refining, evaluating, calibrating, or otherwise improving the detection system.
Substack’s Publisher Agreement, updated July 20, 2026, grants Substack a limited license to submit creator content to automated analysis tools, including AI-based detection systems, for platform safety, content integrity, compliance with Substack’s Terms of Use, and compliance with applicable law.
The Agreement also states that Substack may use automated tools to identify AI-generated or AI-assisted content, detect policy violations, and ensure platform integrity. It states that Substack will provide a Creator an opportunity to dispute a classification result that affects how the Creator’s content is displayed to Readers.
In the Ownership section, the Agreement states that “any third-party tools used for automated analysis are prohibited by contract from retaining or using your content for any purpose other than performing the analysis on our behalf.” In the Automated Content Analysis section, it states that “Automated tools used for this purpose are prohibited from retaining your content or using it for any purpose other than performing the analysis on our behalf.”
Together, these provisions prohibit the tools they cover from retaining “your content” and limit those tools’ use of it to performing the analysis on Substack’s behalf. They do not specify:
how “retaining” is defined or distinguished from temporary storage or processing;
whether performing the analysis involves storage for any period and, if so, for how long;
when the analysis begins and ends;
whether troubleshooting, quality review, error reporting, or dispute resolution is part of the analysis;
what falls within “your content” for purposes of these restrictions;
whether temporary copies and transformed representations are treated as content; or
whether and how these restrictions apply to classifications, logs, metadata, and other data or records derived from or associated with the content.
Pangram’s Data Privacy FAQ states that text submitted through its dashboard and browser extension is stored to display results and maintain a history of queries. It further states that submitted content remains stored for as long as the user’s account is active and can be removed from Pangram’s servers by deleting it from the user’s history.
Pangram’s Privacy Policy states that:
submissions and associated metadata may be collected for registered Customers;
submissions are processed for purposes defined in Pangram’s contract with the Customer;
Pangram personnel may access submission data to diagnose and resolve reported problems;
data may be processed for security, debugging, and maintenance;
submissions are not used to train, develop, refine, or improve AI or machine-learning systems;
submissions are not used to build new products or expand existing products;
content submitted to registered accounts is deleted within thirty days after the Customer account closes, or according to the applicable Customer agreement; and
deleting a query removes the submitted query details while leaving a record that the query occurred for billing and analytics.
Pangram also advertises zero-data-retention options for enterprise API customers. Its public API materials do not state whether the Substack integration uses such an option.
Pangram’s Data Privacy FAQ says Pangram does not share or sell user data to third parties. Pangram’s Privacy Policy separately states that Pangram engages vendors for infrastructure hosting, analytics, authentication, payment processing, customer support, and other services. Those statements may use “share” in different legal or operational senses, but the public materials do not explain which vendors can process information associated with Substack scans.
Pangram’s Terms of Service apply to people and organizations that access or use Pangram’s service. They define submitted text as User Content, reserve Pangram’s ability to monitor information transmitted through the service for operational and other purposes, and grant Pangram a broad license to use feedback submitted by its users.
The terms do not explain whether or how they apply to a Substack publisher who has no Pangram account and whose work is scanned at a reader’s request. They also do not explain whether they apply when a publisher initiates a scan solely because Substack requires an analysis before the publisher can disable detection.
Substack’s Privacy Policy states that Substack may share Personal Information with service providers, including providers of generative-AI services, hosting, maintenance, security, customer support, analytics, content delivery, cloud storage, and cloud computing. The service-provider disclosure does not identify Pangram by name or explain which provider category applies to this integration.
Substack’s Publisher Agreement also states that a Creator is the controller of personal data appearing in the Creator’s published content and that Substack may act as a processor when processing that data on the Creator’s behalf. Its Data Processing Addendum permits Substack to appoint subprocessors subject to applicable contractual protections.
These documents may describe different products, legal relationships, and technical arrangements. I have not located a public, integration-specific notice that reconciles them.
Terms used in this inquiry
For clarity, the following descriptions explain how I use these terms in this inquiry:
Built-in feature means Substack’s reader-facing Scan for AI text feature.
Publisher includes a Substack Creator and any person whose post, Note, comment, or reply is eligible for scanning.
Content means the original submitted writing and any portion, copy, or representation of it.
Derived data means classifications, percentages, confidence values, highlighted passages, segment-level results, embeddings, hashes, fingerprints, logs, identifiers, and other outputs or records created from or associated with the content.
Processing, as used in this inquiry, refers broadly to any operation performed on or involving data, including receiving, collecting, accessing, transmitting, using, storing, retaining, caching, segmenting, transforming, analyzing, reviewing, logging, displaying, disclosing, deleting, or otherwise handling data.
These descriptions are used only to clarify the questions. Any legal definitions or classifications remain governed by the applicable documents and law.
What I am asking
1. Which terms, parties, and privacy roles govern this integration?
Please identify the documents and contractual arrangements that govern content processed through Substack’s built-in feature.
Does a Substack scan fall under:
Pangram’s published Privacy Policy;
Pangram’s Data Privacy FAQ;
Pangram’s Terms of Service;
a separate agreement between Pangram and Substack;
Substack’s Publisher Agreement;
Substack’s Privacy Policy;
Substack’s Data Processing Addendum; or
some combination of these terms?
When a publisher initiates a Pangram analysis solely because Substack requires the analysis before the Disable detection control becomes available, does that action constitute access to or use of Pangram’s service, acceptance of Pangram’s Terms of Service, or the creation of a direct legal relationship between Pangram and the publisher?
If not, which terms govern Pangram’s processing of the publisher’s content, the resulting analysis, associated records, and any feedback or dispute materials?
Where the Substack-specific agreement differs from Pangram’s general policies or terms for direct account-holding customers, does that agreement supplement, modify, or supersede those general terms?
Please identify which published provisions apply to Substack scans and which practices are governed by nonpublic integration-specific terms.
For purposes of Pangram’s Privacy Policy:
Is Substack the “Customer”?
Is the publisher treated as a User, data subject, content owner, third party, or another category?
Is the reader initiating the scan treated as a User?
Does the answer differ when the publisher initiates the scan?
Can Pangram confirm that the restrictions described by Substack apply to every scan conducted through the built-in feature, including publisher-initiated and reader-initiated scans?
Where submitted content or scan-related records contain personal data:
For each category of data and each purpose of processing, is Substack acting as a controller, processor, or both?
Substack currently identifies Pangram in Appendix 3 to the DPA as a subprocessor. Does that designation cover every processing activity associated with Scan for AI text, or only processing performed by Substack on the publisher’s behalf?
Does Pangram act in any additional capacity for any part of the integration, including as a service provider, processor, independent controller, or another category under applicable data-protection law?
Do the roles of Substack or Pangram change depending on whether the scan is initiated by a reader, initiated by a publisher, required before detection can be disabled, initiated by Substack for platform purposes, or reviewed through an error report or dispute?
Which party determines the purposes and means of processing for each activity?
Which party is responsible for responding to privacy requests from the publisher, the person initiating the scan, and any person whose personal data appears in the scanned content?
What transfer mechanisms apply where the data is processed outside the relevant person’s jurisdiction?
Where commercial portions of the Substack-Pangram agreement are confidential, please provide a nonconfidential description of the data-protection terms that govern publisher content.
2. What triggers transmission to Pangram or another analysis provider?
Does Pangram receive Substack content only when a publisher or reader expressly selects Scan for AI text or Check for AI?
If neither the publisher nor any reader initiates a scan, is any portion of the content or associated metadata nevertheless transmitted to, received by, or analyzed by Pangram?
Please distinguish between processing that occurs:
while a draft is created or saved;
before publication;
during publication;
immediately after publication;
when a reader opens a post;
when a reader opens the scan menu;
only after the reader selects the scan command;
through background, batch, or periodic processing;
through precomputation or caching of classifications;
through moderation or policy-enforcement workflows;
through quality-assurance, security, or testing workflows; and
during preparation for a possible dispute or review.
Please answer separately for:
free posts;
paid or subscriber-only posts;
private publications;
Notes;
comments and replies;
unpublished drafts;
scheduled posts;
edited posts;
posts published before July 21, 2026;
content viewed on standalone Substack websites or custom domains;
content delivered by email; and
content viewed through the Substack Reader or app.
Where Substack says the reader-facing feature is unavailable on a particular surface, does “unavailable” mean no Pangram processing occurs, or only that the reader-facing control and result are unavailable there?
If the same item is scanned more than once:
Is the complete content transmitted each time?
Is a previously generated result reused?
Is a cached version of the content or result maintained?
Does editing the content invalidate the previous result?
Apart from Pangram, does Substack submit content to any other third-party automated-analysis provider under the license described in the Publisher Agreement?
For each such provider, please identify:
the provider or provider category;
the purpose of the analysis;
what triggers the submission;
what content and metadata are transmitted;
what is retained or derived; and
what controls are available to the publisher.
3. What information is transmitted during a scan?
Please identify whether Pangram receives:
the complete text or only selected portions;
hidden, deleted, or unpublished portions of a draft;
the title or subtitle;
the publication name;
the publisher’s name, user ID, or account ID;
the content URL or internal content identifier;
the identity or account information of the reader initiating the scan;
the reader’s subscription or access status;
device, browser, session, IP address, or network information;
timestamps;
language information;
revision or version information;
formatting, links, citations, captions, or footnotes;
information identifying whether the content is free, paid, private, or public;
moderation or policy-enforcement flags; or
any other metadata.
For every transmitted category, please state:
which company sends it;
which company receives it;
why it is necessary;
whether it is required to generate the classification;
whether it is retained; and
whether it is disclosed to another service provider or subprocessor.
4. What does “performing the analysis” include?
Please define the beginning and end of “performing the analysis” for purposes of the Publisher Agreement.
During that period, does Pangram, Substack, or any participating service provider:
temporarily store or cache the complete text;
place content in a processing queue;
create copies in working memory;
tokenize, segment, normalize, or otherwise transform the text;
create embeddings, hashes, fingerprints, or numerical representations;
preserve highlighted passages or sentence-level results;
place the text or excerpts in application, diagnostic, security, or error logs;
compare the content with other submissions;
permit personnel to access it for customer support or troubleshooting;
use it for security, abuse prevention, debugging, or maintenance;
use it for quality assurance, benchmarking, calibration, validation, or evaluation;
use it to test a new model or model version;
transmit it to infrastructure, cloud, analytics, security, or customer-support providers; or
process it in connection with a detection-error report or classification dispute?
For each applicable activity, please identify:
its purpose;
its duration;
who can access the information;
where the processing occurs;
whether the activity is considered part of the initial scan;
whether it continues after the result is returned; and
whether any resulting record survives completion of the scan.
Does any party have permission to use content, metadata, classifications, or related records for its own independent purposes, including analytics, product development, service improvement, research, benchmarking, or quality assurance?
5. How are the Publisher Agreement’s restrictions on retention and use implemented?
The Ownership and Automated Content Analysis sections of the Publisher Agreement contain restrictions concerning the retention and use of content by tools used for automated analysis.
Please explain how those provisions operate together, how Substack and Pangram define “retaining,” and whether performing the analysis involves storing, caching, queuing, holding in working memory, or reprocessing content for any period.
If so, please state whether and for how long those activities occur during analysis, troubleshooting, error review, support, security work, or a classification dispute, and identify the term or rule governing each activity.
How do the companies distinguish:
temporary processing from retention;
operational storage from persistent storage;
content from metadata;
content from transformed representations;
content from derived data; and
a completed analysis from an analysis that remains subject to review?
Please state whether “content” includes:
the original text;
excerpts;
tokens or segments;
cached copies;
prompts or model inputs containing the text;
embeddings or numerical representations;
hashes or fingerprints;
highlighted passages;
sentence-level analysis;
revision information; and
material resubmitted or preserved for a dispute.
6. What retention configuration applies to Substack scans?
Substack’s Publisher Agreement states that automated tools used for this purpose are prohibited from retaining “your content.”
Pangram’s public materials describe different data-handling arrangements for its products. Its Data Privacy FAQ states that text submitted through its dashboard and browser extension is stored to display results and maintain query history, and that submitted content remains stored for as long as the user’s account is active unless deleted from that history. Its API page separately advertises “zero data retention options” for enterprise customers.
Those materials do not identify the configuration used for the Substack integration or explain:
whether Pangram’s advertised zero-data-retention option is used for Substack scans;
whether a separate integration-specific configuration applies; or
whether “zero data retention” is equivalent in scope to the Publisher Agreement’s prohibition on retaining “your content.”
Please identify the technical and contractual retention configuration used for Substack scans. Does the integration use:
Pangram’s enterprise zero-data-retention option;
a separate integration-specific configuration;
a configuration under which “your content” is not retained but classifications, metadata, logs, or other related records may be retained; or
another arrangement?
If any aspect of Pangram’s account-history model applies, please identify what is stored, for how long, and how that practice is reconciled with the Publisher Agreement.
Please identify what happens to each of the following after a result is returned:
the original content;
excerpts and transformed representations;
the classification percentage;
the overall label;
confidence values;
segment-level classifications;
highlighted passages;
scan timestamps;
publisher and publication identifiers;
reader identifiers;
content URLs or internal IDs;
diagnostic, security, and audit logs;
records that a scan occurred;
billing or analytics records;
detection-error reports;
dispute materials; and
records associating a result with a particular work or publisher.
For each retained category, please state:
which company retains it;
its retention period;
its purpose;
the contractual or policy basis for retention;
who can access it;
whether it appears in a dashboard or query history;
whether it is stored in backups or disaster-recovery systems; and
how the publisher can request access to or deletion of it.
Is any dashboard or query-history entry created in a Pangram account, workspace, or other environment used for the Substack integration?
If so:
Can Substack personnel view it?
Can Pangram personnel view it?
Is it automatically deleted?
Does the entry include the original text or only the result?
Does the Pangram practice of retaining a record that a query occurred for billing and analytics apply to Substack scans? If so, what information is contained in that record?
What happens to relevant records when:
detection is disabled;
a post or Note is edited;
a post or Note is deleted;
a publication is deleted;
a publisher leaves Substack;
a reader deletes their account;
Substack terminates its Pangram relationship; or
a legal hold or anticipated dispute arises?
Does Substack audit, test, or otherwise verify Pangram’s compliance with the applicable retention and use restrictions?
7. Which service providers and subprocessors participate?
Substack’s Privacy Policy permits service-provider processing across several categories. Pangram’s Privacy Policy identifies vendors involved in infrastructure hosting, analytics, authentication, payment processing, customer support, and other services.
For the built-in feature, please identify which categories of Substack and Pangram providers receive, transmit, store, or can access:
publisher content;
transformed representations of content;
publisher, publication, reader, URL, and timestamp metadata;
classification results;
highlighted passages;
technical, diagnostic, security, and audit logs;
detection-error reports; and
classification-dispute materials.
Please clarify:
whether Pangram is treated as a Substack service provider or subprocessor;
whether Pangram’s own vendors are subprocessors for purposes of this integration;
whether the Publisher Agreement’s restrictions flow down to every participating vendor;
whether every provider is prohibited from using the information for its own purposes;
whether any provider may use the information for analytics, product development, benchmarking, service improvement, or research;
what countries or regions the information may be processed in;
whether and by what means Substack audits, tests, monitors, or otherwise verifies the compliance of Pangram and any participating subprocessors with the applicable contractual, technical, privacy, security, retention, deletion, and use restrictions;
how frequently such verification occurs, whether it includes downstream vendors and technical controls governing retention and deletion, and what remediation or enforcement measures apply when noncompliance is identified; and
which company remains accountable if a downstream provider violates the applicable restrictions.
Pangram’s Data Privacy FAQ says Pangram does not share data with third parties, while its Privacy Policy says third-party vendors provide infrastructure, analytics, authentication, customer support, and other services.
Please explain how Pangram defines “share” in the FAQ and whether any Pangram vendor processes content or related data from Substack scans.
8. When was Pangram appointed, what notice was provided, and what objection rights apply?
Substack currently identifies Pangram in Appendix 3 to the DPA, which supplies the Annex III list of subprocessors.
Appointment and notice
Please state:
the date Substack appointed Pangram as a subprocessor or otherwise authorized Pangram to receive or process Creator content or personal data through the integration;
the date Substack and Pangram entered into the applicable data-processing arrangement;
the date Pangram first became technically capable of receiving or processing content through the integration;
the date Pangram first actually received or processed Creator content or personal data through the integration;
the date the subprocessor list was updated to include Pangram and the date that updated list was first made publicly available, if different;
the date and method by which notice of Pangram’s addition was provided;
whether notice was provided to all affected Creators or only to Creators who had previously requested email notifications through the address identified in the Publisher Agreement;
what form of notice was provided to Creators who had not requested email notifications;
whether the July 20, 2026 update to the Publisher Agreement was intended to provide notice of Pangram’s appointment; and
where a copy or archived version of each applicable notice may be reviewed.
Which notice and objection periods applied?
The Publisher Agreement and incorporated data-protection terms describe two advance-notice periods and one post-update objection period:
at least six weeks’ prior notice of changes to the subprocessor list under the DPA;
twenty-one calendar days’ notice of intended additions or replacements under the incorporated Standard Contractual Clauses, where those clauses apply; and
thirty days following the update of “this page,” during which a Creator may submit a written objection.
Please explain:
which of these provisions applied to Pangram’s appointment;
which Creators were entitled to each form of notice or objection;
what event triggered each applicable period;
the date on which each applicable period began;
the date on which each applicable period ended or will end;
whether any applicable period remains open;
whether Pangram began receiving or processing Creator content before any applicable notice or objection period expired and, if so, on what basis;
how the six-week and twenty-one-day advance-notice provisions interact;
how the thirty-day written-objection provision interacts with the DPA’s statement that a Creator may object by exercising termination rights;
whether a written objection submitted during the thirty-day period triggers any review, consultation, alternative arrangement, suspension, or other process apart from termination;
which provision governs where the periods or procedures differ; and
whether Substack treats the appointment as accepted before the applicable objection period has expired.
What objection right is available to a Creator established in the United States?
The main Privacy section of the Publisher Agreement and the opening language of Annex 1 describe the DPA’s territorial scope in somewhat different terms. The main section includes Creators based in the EEA, UK, or Switzerland, as well as certain processing concerning data subjects in those locations. Annex 1 refers to Creators based in the EEA or UK and processing governed by the GDPR or its UK equivalent.
As a Creator established in the United States, please clarify:
whether I may exercise the written subprocessor-objection right described in the main body of the Publisher Agreement;
whether that right is available to every Creator or only to a Creator whose processing falls within the scope of the DPA;
whether the DPA and its objection procedure apply to a US-based Creator whose publication includes subscribers, readers, authors, or other data subjects located in the EEA, UK, or Switzerland;
how a US-based Creator can determine whether the DPA applies to their publication and to Pangram’s processing;
what procedure applies to an objection;
the email or postal address to which an objection should be submitted;
the information an objection must contain;
the deadline that applies;
whether submitting an objection suspends Pangram’s processing of the objecting Creator’s content while the objection is considered;
whether an objection can prevent Pangram from processing that Creator’s content;
whether Substack will provide an alternative technical or provider arrangement;
whether the objection triggers consultation or another resolution process;
whether the sole available consequence is termination of the Publisher Agreement; and
if the objection right is unavailable to a US-based Creator, what mechanism, if any, allows that Creator to object to or prevent processing by a newly appointed subprocessor.
Please also state whether a Creator’s continued use of Substack during an unresolved objection is treated as acceptance of the subprocessor.
9. Can a publisher prevent analysis before content is transmitted?
Does Substack provide a control that allows a publisher to prevent content from being submitted to Pangram before any scan occurs?
Can a publisher disable processing:
at the account level;
at the publication level;
for all posts;
for all Notes;
for comments and replies;
for paid or private content;
for a particular item;
before publication; and
before a draft is submitted for an initial analysis?
Can a person whose comment or reply is eligible for scanning disable detection on that text, or is that control available only to publication owners?
Does the Block AI training setting have any effect on the Substack-Pangram integration, or does it apply only to external AI crawlers?
Does disabling the reader-facing feature prevent other automated analysis conducted for platform safety, content integrity, moderation, policy enforcement, or legal compliance?
10. Can a publisher decline or limit the automated-analysis license?
The Publisher Agreement grants Substack a license to submit creator content to automated analysis tools.
Can a publisher decline or limit Substack’s exercise of that license:
at the account level;
at the publication level;
for an individual post or Note;
for comments or replies;
for paid or private content; or
before content is submitted to an external provider?
Does selecting Disable detection alter, limit, or withdraw the license, or does it affect only the feature’s operation or reader-facing display?
If Substack provides no way to limit this portion of the license, please confirm whether the only way for a publisher to prevent the license from applying to a particular work is to refrain from publishing that work on Substack.
11. Does the automated-analysis license extend to additional, future, or replacement providers and purposes?
The Publisher Agreement grants Substack a license to submit Creator content to automated analysis tools, including AI-based detection systems, for the stated purposes of “platform safety,” “content integrity,” compliance with the Terms of Use, and applicable law. The provision does not identify Pangram by name and is not limited on its face to a single provider.
Please clarify:
whether this license permits Substack to submit Creator content to additional, future, or replacement providers without obtaining further agreement from the Creator;
whether the license also covers automated-analysis tools operated directly by Substack;
whether the appointment, addition, or replacement of a provider requires:
an amendment to the Publisher Agreement;
notice under the Changes to this Agreement provision;
an update to Appendix 3 or Annex III;
another form of notice; or
no additional notice;
whether a provider may process Creator content under this license without being listed as a subprocessor, including where Substack characterizes that provider in another legal role or determines that the processing does not involve personal data;
what advance notice, objection right, or opt-out opportunity applies before a new or additional provider begins processing Creator content;
whether the contractual restrictions on retention and use beyond performing the analysis bind every additional, future, or replacement provider and each of its service providers or subprocessors;
whether Substack audits or otherwise verifies those restrictions before a new provider begins processing;
whether a Creator’s existing settings, including Disable detection on individual posts or Notes, automatically carry forward to a new provider;
whether a new or replacement provider may rescan previously published content, earlier versions of edited content, or content for which detection was previously disabled;
whether a material change to a provider’s model, technical architecture, data flow, retention configuration, service providers, or subprocessors triggers notice to Creators; and
which forms of automated analysis Substack currently conducts within each of the stated purposes, and which internal or third-party tools perform them.
Please also clarify whether the license authorizes automated analysis beyond reader-facing AI detection, including analysis for policy enforcement, moderation, legal compliance, account review, distribution, security, or other platform purposes, and identify the controls available to a Creator for each form of processing.
12. Why must a publisher generate an analysis before disabling detection?
Substack’s instructions require a publisher to submit a draft post or Note for Pangram analysis before the Disable detection control becomes available.
Please explain the purpose of requiring this initial analysis.
In particular:
Why is a completed Pangram analysis technically, operationally, or contractually necessary before detection can be disabled?
Is the prerequisite essential to the operation of the feature, or is it a product-interface choice?
Could Substack provide a pre-scan control that allows a publisher to decline analysis before any content is transmitted to Pangram?
For whose benefit and for what purposes is the required analysis performed?
Which parties may access, receive, retain, rely on, or otherwise use the analysis, its result, or any associated records?
Who can access the resulting analysis before and after the publisher selects Disable detection?
Is the result used solely to present the publisher with the report from which the control is accessed?
Is the result also used for moderation, platform integrity, policy enforcement, quality assurance, analytics, dispute preparation, or another purpose?
If the publisher does not select Disable detection, does the initial publisher-generated analysis become the result later shown to readers, or do reader requests generate separate analyses?
Does the initial analysis create a cached classification, query-history entry, internal record, or other reusable result?
If the publisher selects Disable detection, are the analysis and all associated original, transformed, and derived records deleted, or are any of them preserved?
Can the publisher separately request deletion of the initial analysis and associated records?
Does generating the analysis have any effect on the treatment, ranking, distribution, moderation, monetization, or internal assessment of the content, regardless of whether detection is later disabled?
Why does exercising the control require the publisher first to initiate the processing that the publisher may be seeking to prevent?
Please explain whether Substack plans to make the control available before any transmission or analysis occurs.
13. What does Disable detection actually disable?
Does disabling detection on a post or Note prevent that content from being submitted to Pangram for future analysis?
Or does the content remain eligible for submission and analysis while the result is withheld from readers?
Does disabling detection:
prevent all future scans;
suppress only the displayed result;
delete an existing result;
delete the initial scan required to access the control;
delete any Pangram history or dashboard record;
affect classifications generated before the setting was selected;
affect processing for moderation or platform integrity;
persist after the post or Note is edited; and
apply to all versions of the content?
Substack’s instructions require a publisher to generate a Pangram analysis before the Disable detection control becomes available.
During that initial scan:
Is the complete draft transmitted to Pangram?
What metadata accompanies it?
What records does Substack retain?
What records does Pangram retain?
Does selecting Disable detection subsequently delete any of those records?
A publisher seeking to decline external analysis should have a way to exercise that choice before the content is submitted for analysis. Does Substack plan to provide such a control?
14. What notice, visibility, and downstream effects follow a scan?
Is the publisher notified when their content is scanned?
Can the publisher see:
whether a scan occurred;
when it occurred;
how many scans occurred;
whether the scan was initiated by the publisher or a reader;
which version of the content was scanned;
the result shown to the reader; and
whether the result has changed over time?
If revealing the reader’s identity would implicate that reader’s privacy, can Substack provide aggregate or non-identifying scan information to the publisher?
Who can see the classification:
only the person who initiated the scan;
all subscribers;
all readers with access to the item;
the publisher;
Substack personnel;
Pangram personnel; or
other parties?
How long does the reader-facing result remain available?
Is it generated anew for each reader, or is one result cached and displayed repeatedly?
Is the classification used for any purpose beyond responding to the scan request, including:
recommendations;
search or discovery ranking;
distribution;
email delivery;
monetization eligibility;
advertising eligibility;
editorial selection;
moderation;
policy enforcement;
account review;
fraud or abuse detection;
trust or reputation scores; or
internal research and analytics?
If a Pangram result contributes to any consequential decision affecting reach, revenue, account status, or content availability:
Is human review required?
Is the publisher notified?
Can the publisher inspect the evidence?
Can the publisher appeal?
Is the Pangram score treated as one signal or as a determinative classification?
15. What records does each company maintain, and what rights does the publisher have?
Please identify the records maintained by each company in connection with:
the scan request;
the content submitted;
the person initiating the scan;
the classification;
the displayed result;
the version of the work analyzed;
any internal use of the result;
detection-error feedback;
support communications; and
a classification dispute.
Can a publisher request:
confirmation that their work was processed;
a copy of the information transmitted;
a copy of the classification and segment-level analysis;
access to associated metadata;
correction of inaccurate records;
deletion of the original content;
deletion of derived records;
restriction of further processing; and
an accounting of the providers that received the information?
Does that pathway remain available when the publisher:
has no Pangram account;
did not initiate the scan;
cannot view Pangram’s query history;
is the author of a comment or reply rather than the publication owner; or
has deleted the work or left Substack?
If requests must be submitted through Substack, please identify the process. If Pangram accepts direct requests, please explain how it verifies that the requesting person is the publisher or author of the scanned work.
16. How does the opportunity to dispute a result described in the Publisher Agreement work?
Does Report detection error constitute the dispute opportunity mentioned in the “Automated Content Analysis” section of the Publisher Agreement?
If not, where and how can a publisher dispute a result?
Please clarify:
what question the dispute process is intended to decide, including whether it reviews the technical reliability of the classification, the appropriateness of continuing to display it, or the publisher’s authorship itself;
whether the disputed result receives human review;
which company conducts the review;
the qualifications, role, and decision-making authority of the reviewer;
whether Pangram reruns the analysis;
whether the same model and model version used for the original classification are preserved and reviewed;
whether a different model or model version is used during the review;
how conflicting results between the original scan and a later scan are resolved;
what standard or threshold governs whether a classification is corrected, withdrawn, or removed;
which party bears responsibility for supporting the continued display of the disputed classification;
whether a publisher can obtain a complete review based on the disputed classification and the companies’ own records without disclosing unpublished drafts, revision histories, notes, research materials, source files, or details of the publisher’s creative process;
whether preserving the confidentiality of those materials carries any adverse inference or limits the correction, withdrawal, appeal, or other remedy available to the publisher;
whether an inaccurate or insufficiently supported result can be corrected, withdrawn, or removed;
whether the publisher receives a written decision;
whether the decision explains the information reviewed, the applicable standard, the findings reached, and the basis for the outcome;
whether an appeal or escalation process exists;
what happens to the reader-facing classification while review is pending;
whether the classification is marked as disputed during the review;
whether cached results are updated following a successful dispute;
whether readers who previously viewed the classification receive any correction or notice;
what response timeframe applies; and
whether repeated erroneous classifications concerning the same publisher, publication, or type of writing can be escalated for broader review.
What information is reviewed during a dispute?
Please identify the information available to the reviewer, including:
the exact version of the work originally scanned;
the original model and model version;
the original classification and percentage;
confidence values;
segment-level analyses;
highlighted passages;
the time and circumstances of the scan;
the identity or category of the person who initiated it;
system, diagnostic, and audit records;
known limitations or error conditions relevant to the classification;
prior scans of the same work; and
any subsequent scans or conflicting results.
If the original content is not retained, what information is reviewed when deciding the dispute?
Is the publisher required to resubmit the published work?
Does the review rely on:
a cached copy;
excerpts;
tokens or segments;
embeddings or hashes;
highlighted passages;
the original result;
diagnostic logs;
a newly submitted copy of the work; or
another preserved representation?
Which company retains each dispute-related record, for how long, and under which contractual or policy provision?
What protections apply to nonpublic materials?
A publisher should not have to disclose confidential source materials or details of their creative process merely to obtain review of a classification generated and displayed through Substack’s systems.
If Substack or Pangram may request, invite, or accept unpublished drafts, revision histories, research, notes, source files, correspondence, metadata, or other nonpublic materials during a dispute, please clarify:
whether the submission of those materials is entirely optional;
whether a complete review remains available without them;
whether declining to provide them carries any adverse inference;
whether narrower or redacted materials may be submitted instead;
what written confidentiality and data-processing terms apply before submission;
whether the publisher receives those terms before deciding whether to submit anything;
which company receives the materials;
which personnel, contractors, service providers, or subprocessors may access them;
whether the materials are used solely to resolve the individual dispute;
whether they may be used for training, model development, fine-tuning, evaluation, calibration, benchmarking, validation, quality assurance, product improvement, research, or any other secondary purpose;
whether they are associated with the publisher’s account, publication, or other works;
how they are secured during transmission and storage;
how long they and any copies, excerpts, summaries, or derived records are retained;
whether they remain in backups, logs, or disaster-recovery systems;
when and how they are deleted after the dispute;
whether the publisher may withdraw the materials or request their deletion before the dispute concludes;
whether the publisher receives confirmation of deletion;
whether the materials may be disclosed in response to legal process and what notice would be given;
whether the publisher may designate an authorized representative to participate in the dispute; and
what remedy is available if the materials are used, retained, or disclosed outside the stated terms.
Does the dispute process determine whether Substack has sufficient grounds to continue displaying its classification, or does it purport to determine whether the publisher created the work?
17. How is Report detection error feedback used?
Substack says error feedback will be reviewed to improve detection quality. Pangram’s Privacy Policy says submissions are not used to train, develop, refine, or improve AI or machine-learning systems.
Please reconcile those statements.
When a person submits a detection-error report:
Does Substack send the report to Pangram?
Does the report include the complete content?
Does it include highlighted passages, classifications, metadata, or account identifiers?
Is the content scanned again?
Is the feedback used only to resolve the individual report?
Is it used for model evaluation, benchmarking, calibration, validation, quality assurance, rule changes, or future product development?
Is it added to a research, testing, or evaluation dataset?
Is any portion used to train or fine-tune a model?
How long is the report retained?
Can the person who submitted it delete or withdraw it?
Pangram’s Terms of Service grant Pangram broad rights in “Feedback” submitted by Pangram users.
Is feedback submitted through Substack’s Report detection error process treated as Feedback under Pangram’s Terms of Service?
If so:
On what basis are those terms applied to a publisher who has no Pangram account?
What rights does Pangram receive?
Does the license apply only to the publisher’s comments about the service, or also to attached text, drafts, evidence, and other materials?
18. What happens to scan-related records when content, a publication, or an account is deleted?
The DPA addresses deletion or return of personal data processed on the Creator’s behalf upon termination of the Publisher Agreement. It does not expressly map that obligation onto scan-related content, classifications, metadata, transformed representations, or other derived records, and it does not explain what happens when an individual item or publication is deleted without termination of the entire Publisher Agreement. The Agreement also allows certain irreversibly anonymized aggregate outputs to be retained outside the DPA’s deletion or return obligations.
Please explain what occurs when:
a post, Note, comment, or reply is deleted;
a post or Note is unpublished;
a publication is deleted;
a Creator account is deleted;
the Publisher Agreement is terminated; or
a Creator separately requests deletion of scan-related information.
For each event, please state what happens at Substack, Pangram, and every participating service provider or subprocessor to:
the version of the work that was scanned;
excerpts or cached copies;
tokenized, segmented, embedded, hashed, fingerprinted, or otherwise transformed representations of the work;
the overall classification and percentage;
confidence values;
segment-level analyses and highlighted passages;
records that a scan occurred;
timestamps, URLs, content identifiers, and publication identifiers;
information identifying or describing the person who initiated the scan;
query-history or dashboard entries;
detection-error reports and any information associated with them;
dispute records, correspondence, and decisions;
technical, diagnostic, security, billing, analytics, or audit logs;
backups and disaster-recovery copies;
research, statistical, benchmarking, or quality-assurance records; and
anonymized, de-identified, aggregated, or other derived data associated with the scan.
For each category, please state:
whether it is deleted, retained, anonymized, de-identified, or aggregated;
which company retains or deletes it;
the time required for deletion to propagate through each company and participating provider;
the retention period where it remains;
the purpose and contractual or policy basis for continued retention;
whether it remains linkable to the Creator, publication, work, URL, or person who initiated the scan;
whether it remains in active systems, archives, backups, logs, or disaster-recovery systems;
whether legal holds, complaints, anticipated disputes, security investigations, or other exceptions extend the retention period;
whether the Creator receives notice of any such exception;
whether deletion by Substack automatically causes Pangram and downstream providers to delete corresponding records;
whether the Creator must submit a separate request to Pangram or another provider; and
whether Substack provides confirmation when deletion has been completed throughout the processing chain.
Please also clarify:
whether deletion of the underlying work removes any cached or reader-facing classification;
whether deletion prevents the result from being reused in future scans, feed-health calculations, statistics, research, benchmarking, or product analysis;
whether a Creator may request deletion of scan-related records without deleting the underlying work, publication, or account;
whether disabling detection has any deletion effect distinct from deleting the work;
whether an anonymized or aggregated record may survive after the original work and direct identifiers are deleted; and
what standard Substack uses to determine that surviving data can no longer be associated, directly or indirectly, with the Creator, publication, or work.
19. Which company is responsible for support and enforcement?
Which company is responsible for answering questions or resolving concerns involving:
data transmission;
temporary processing;
retention;
access and deletion;
classifications;
technical errors;
disputes;
detection-error feedback;
privacy rights;
downstream service providers; and
enforcement of the Substack-Pangram contractual restrictions?
Please provide:
a public contact for integration-related questions;
a privacy contact;
a dispute contact;
a support-ticket process;
an escalation route;
expected response timeframes; and
a process for publishers who have no Pangram account.
Please clarify whether a publisher can contact Pangram directly about the processing of their work even when the publisher did not initiate the scan.
If Pangram must refer the publisher to Substack, please explain how Substack can obtain and provide the records necessary to answer the publisher’s question.
20. Will the companies publish an integration-specific data-practices notice?
Will Substack and Pangram publish a concise, integration-specific notice covering:
the data-flow sequence;
the events that trigger transmission;
all eligible content types and surfaces;
the information transmitted;
the roles of each company;
the governing terms;
participating service providers and subprocessors;
the meaning of “performing the analysis”;
the applicable retention model;
the treatment of temporary copies and derived data;
the treatment of repeated scans;
publisher controls;
privacy requests;
dispute procedures;
downstream uses of classifications; and
support contacts?
A data-flow diagram showing what happens before, during, and after a scan would be particularly useful.
Please also provide a dated change log and advance notice when material data practices, providers, retention arrangements, or publisher controls change.
A related but distinct set of questions about Pangram’s browser extension
This section concerns Pangram’s browser extension and any extension-based scanning that may occur separately from Substack’s built-in feature.
Pangram states that its browser extension is available on Substack, labels content users browse, and actively scans portions of Substack and other platforms. Pangram also states that extension data is stored by default for the user’s Feed Health Screenshot, while users may change their storage preferences. It separately encourages users to opt in to data collection for research into the proliferation of AI content.
21. How does Pangram’s browser-extension scanning affect Substack publishers?
What is the relationship between Substack and Pangram’s browser extension?
Please describe any technical, commercial, or contractual relationship between Substack and Pangram concerning the browser extension.
In particular:
Does Substack authorize, facilitate, promote, technically support, or otherwise coordinate with Pangram concerning extension-based scanning of Substack content?
Does Substack provide Pangram with an API, feed, privileged access, content identifiers, authentication mechanism, or other technical assistance for that activity?
Does Substack receive classifications, statistics, research findings, analytics, or other information generated through the extension?
Does any agreement between Substack and Pangram govern the extension?
Does Substack impose any contractual or technical restrictions on Pangram’s extension-based access to Substack content?
If the extension operates without any Substack integration, privileged access, data transmission, coordination, or contractual arrangement, please state that clearly.
How does the extension obtain and scan Substack content?
When a person using the Pangram browser extension visits a Substack publication:
Is eligible text automatically submitted to Pangram as the person browses or scrolls, or only after the person expressly initiates an individual scan?
Does the extension extract text from the page rendered in the user’s browser?
Does it transmit a URL or other identifier that causes Pangram’s servers to retrieve the content directly from Substack?
Does it obtain content through an API, feed, or other mechanism?
Does it scan only content that the user has actively opened and is authorized to view?
Can it independently discover, request, preload, or retrieve additional Substack content that the user has not actively opened?
Is the complete post submitted, or only the portions visible or rendered in the browser?
Are Notes, comments, and replies scanned?
Can the extension scan paid, subscriber-only, private, or otherwise access-controlled content displayed to an authenticated user?
For access-controlled content, please explain the technical and contractual basis for the extension’s processing and the notice, retention, access, deletion, and publisher-control provisions that apply.
What content and metadata are transmitted?
Please identify whether extension-based scanning transmits or otherwise makes available to Pangram:
the complete text or selected portions;
titles and subtitles;
Notes, comments, and replies;
URLs or internal content identifiers;
author or publisher names and identifiers;
publication names;
timestamps;
language information;
subscription or access status;
information identifying the extension user;
Pangram account or feed-history information;
browser, device, session, IP address, or network information; or
any other associated metadata.
Please also clarify:
whether the scanned content is associated with the extension user’s Pangram account, feed history, or Feed Health Screenshot;
whether multiple scans of the same work are linked or aggregated;
whether the publisher, publication, or work can be identified from the stored records; and
whether Substack receives any of this information or any results derived from it.
What is stored, for how long, and for what purposes?
Please identify what Pangram stores by default when its extension scans Substack content, including:
the original text;
excerpts or visible portions;
transformed or segmented representations;
titles, URLs, author names, and publication names;
classification results and percentages;
highlighted passages;
scan timestamps;
feed-history records;
technical, diagnostic, security, or audit logs; and
records associating the scan with a particular user, publisher, publication, or work.
For each category, please state:
the retention period;
the stated purpose;
who may access it;
whether it appears in the extension user’s history or Feed Health Screenshot;
whether it remains in backups, logs, or disaster-recovery systems; and
what event causes it to be deleted.
Pangram states that users may change their storage preferences. Please explain:
which preference controls extension-related storage;
what changing that preference prevents prospectively;
whether it deletes content and records collected previously;
whether it deletes the original text, classifications, metadata, and derived records;
whether any operational, billing, analytics, security, or audit records remain; and
whether the user receives confirmation that deletion has occurred.
How are extension data and research participation separated?
Pangram separately encourages extension users to opt in to data collection for research concerning the proliferation of AI content.
Please clarify:
what information the research opt-in authorizes Pangram to collect or use;
whether the opt-in covers the scanned writing itself, excerpts, classifications, URLs, author or publication identifiers, feed histories, or aggregate statistics;
whether the publisher’s identity or publication can be inferred from the research data;
whether research data is anonymized, de-identified, aggregated, or pseudonymized;
whether extension-scanned content is used for statistics, benchmarking, model evaluation, validation, calibration, quality assurance, product development, or other research;
whether it is used to train, fine-tune, develop, refine, or improve any model or AI system;
whether research participation affects the retention period;
whether withdrawing from research prevents future use of previously collected data;
whether withdrawal causes previously collected research data to be deleted; and
whether opting out of research also prevents operational storage for feed history or the Feed Health Screenshot.
What notice, access, deletion, and opt-out rights does the publisher have?
When a publisher’s work is scanned through Pangram’s browser extension:
Is the publisher or author given notice that the scan occurred?
Can the publisher determine whether, when, or how often the work was scanned?
Can the publisher request confirmation that Pangram processed the work?
Can the publisher obtain access to the submitted text, associated metadata, classification, and other records?
Can the publisher request correction, restriction, or deletion of those records?
Can those rights be exercised when the publisher has no Pangram account and did not initiate the scan?
What process verifies that the requesting person is the publisher or author of the scanned work?
Which company is responsible for responding to such requests?
Is there a publisher-level opt-out from extension-based scanning?
Does Substack provide a setting, technical signal, page-level instruction, or machine-readable preference that Pangram’s extension is expected to honor?
Can the author of a Note, comment, or reply exercise these rights independently of the publication owner?
Do Substack’s publisher controls affect extension-based scanning?
Please clarify whether selecting Disable detection affects Pangram’s browser extension or applies only to Substack’s built-in Scan for AI text feature.
Please also clarify whether Substack’s Block AI training setting:
prevents extension-based scanning;
limits transmission to Pangram;
restricts storage or research use;
communicates a publisher preference that Pangram is expected to honor; or
applies only to external crawlers and model-training activities.
If neither control affects browser-extension scanning, please state that clearly and identify any control available to publishers who wish to prevent or limit that processing.
Finally, please distinguish clearly among:
Substack’s built-in Scan for AI text feature;
any background or platform-directed automated analysis arranged by Substack; and
scanning performed through Pangram’s browser extension or other Pangram products that may operate separately from Substack’s built-in feature.
For each system, please identify who initiates the processing, how the content is obtained, which terms govern it, what publisher controls apply, and which company is responsible for responding to questions or requests.
Why this matters
Publishing makes a work available for reading. Automated analysis and classification arranged through the platform constitute a separate processing activity.
The feature produces a classification in connection with a publisher’s work and may create related data and records. Those outputs may affect reputation, contracts, licensing, editorial and publication decisions, and economic opportunities.
To make informed decisions about whether and how to publish my work on Substack, I need to know what is transmitted, who receives it, for what purposes it is used, what is retained, what controls are available to me, how any classification can be disputed, corrected, or withdrawn, and how related records can be accessed or deleted.
Requested response
I would appreciate a response addressing the numbered questions above.
Where a question is already answered in a publicly available document, please identify the specific section or provision that supplies the answer.
Where an answer depends on the nonpublic agreement between Substack and Pangram, please provide a nonconfidential description of the terms governing data protection, retention, use, and accountability.
Where an answer requires information or confirmation from Pangram, please obtain and include it in Substack’s response.
Where an answer is not currently documented, please provide the answer and update the relevant public documentation.
I would also appreciate a response published at a stable public URL.
This inquiry reflects the public documentation available to me and reviewed on July 23, 2026.
I am sending it to privacy@substackinc.com and publishing it openly.
___
Materials reviewed
All materials were accessed on July 23, 2026.
Substack Help Center, “How can I detect AI on Substack?”, updated July 23, 2026.
Substack, Publisher Agreement, updated July 20, 2026.
Substack, Privacy Policy, updated May 14, 2026.
Substack, Terms of Use, effective April 21, 2025.
Pangram Labs, Privacy Policy, updated August 14, 2025.
Pangram Labs, Data Privacy FAQ.
Pangram Labs, Terms of Service, updated August 14, 2025.
Pangram Labs, AI Detection API.
Pangram Labs, How AI Detection Works.
Pangram Labs, Feed Scanner - Browser Extension.

