Skip to main content

Last updated : 2026-05-29

sen-ai.fr measures how generative AI systems (ChatGPT, Gemini, Claude, Mistral) reference brands when answering real buyer questions. This page documents how we do it - the data sources, the AI providers, the statistical paradigm, and the safeguards we wire in so the numbers are honest. If you're a procurement / DPO / compliance reviewer, this is the page to attach to your DPA. If you're a marketer, this is the page that explains why a single scan isn't enough.

1. The problem we measure

When a user asks ChatGPT "what's a good moisturiser for sensitive skin", the answer surfaces certain brands and not others. That ranking is the new SEO. Unlike Google's search results, AI answers vary across runs - the same question on the same model can produce different brand mixes 30 seconds apart. Anyone measuring AI visibility from one scan is reading noise.

sen-ai.fr's job is to stabilise that signal at scale, then surface the actionable deltas : which questions cite you, which cite competitors, which sources do the AI systems consult, and where on your site do the AI systems land.

2. The N-runs paradigm

Every scan repeats each question N times per provider instead of asking once. Default for paid plans is N=10 ; teams can configure lower for cost or higher for confidence. The output is then averaged with a published variance band so the user knows whether a 22% mention rate is "solid at ±3pts" or "noisy at ±18pts".

Why N=10 specifically : empirical variance for the providers we use is ~3-5% for OpenAI grounded mode, ~10-15% for Gemini ungrounded, ~3-5% for Claude grounded. At N=10 the standard error drops to ±1-2pts across the mix - tight enough to detect quarterly trend changes.

Source for the variance characterisation : the SparkToro study (Fishkin / O'Donnell, January 2026), which measured brand-recommendation inconsistency across 2,961 runs, plus our own internal benchmarks across 50+ scans on French cosmetics, pharma, automotive, and consumer-services verticals.

3. AI providers & models

sen-ai.fr queries third-party general-purpose AI systems in read-only mode. We don't train, fine-tune, or feed responses back into any model. Current mix :

  • OpenAI ChatGPT - GPT-4.1 mini for scan runs. Grounded mode (web search) when available.
  • Google Gemini - Gemini 3.5 Flash for scan runs (the model behind the consumer Gemini app and AI Mode) ; Gemini 3.1 Flash-Lite powers our brand-extraction analyzer. EU-hosted via Google Ireland Ltd.
  • Anthropic Claude - Haiku 4.5 for structured-JSON extraction (brand mention parser, sentiment judge) ; Sonnet 4.6 for editorial synthesis.
  • Mistral Le Chat - planned ; not part of live scans yet. Critical for the French market (41% of Mistral traffic = France).

The mix evolves : providers get added when their adoption crosses the 5% search-share threshold per country. We publish the exact model versions each scan ran on in the per-scan compliance report so historical comparisons stay reproducible.

Model eras. AI vendors replace their models regularly, and a model change can move your visibility numbers without anything changing on your side. Every scan records the model versions it actually ran on. When that mix changes between two scans of the same tracker, your trend is annotated with an era boundary : the data point is marked ("AI models updated") and we never compute a change indicator across the boundary - comparing two different models as if they were the same measurement instrument would produce a misleading curve. The boundary itself is information : it shows how the new model generation sees your brand.

4. Personas & questions

For each tracked brand we generate a topic taxonomy (from the site's ranking keywords on Google), then a set of personas (synthetic user archetypes) anchored on each topic, then 5 question types per persona (informational, commercial, transactional, etc.).

The questions are brand-agnostic on purpose - they describe a buyer situation, not the brand. This avoids leading the AI and matches how real users phrase things in ChatGPT / Perplexity.

The persona + question generators are themselves Claude Haiku calls bound by a strict JSON schema. No human-in-the-loop required for v1 ; the workspace owner can edit any persona or question before launching the scan.

5. Brand classification & mention extraction

Every LLM response is parsed by an "analyser" pass (Claude Haiku, strict JSON output) that extracts every brand-like entity, classifies its sentiment (positive / negative / neutral), and tags whether it's the focus brand or a competitor.

Brand classification is workspace-controlled. The system surfaces every entity it finds ; the workspace owner promotes them to my_brand / competitor / ignored via the Brands tab. Auto-classification is conservative - unknown entities default to discovered and stay invisible from the metrics until the owner curates them. No hardcoded vertical / region / brand lists.

A Haiku-as-judge layer (live) re-reads each negative mention and overturns obvious sentiment false positives (e.g. "not suitable for X" misclassified as negative when the underlying statement is factual). Isolated negative signals are also tempered via conservative severity buckets.

6. Citation extraction

When an AI provider exposes citation metadata (OpenAI grounded mode, Gemini Vertex AI search grounding, etc.), we persist the URLs verbatim. When it doesn't, we parse the response text for explicit URL patterns. Each citation is annotated with the publishing domain, the URL, and a 200-character context snippet.

Citations feed several downstream features : Page Audit (which of YOUR pages get cited), Competitors (which pages of RIVALS get cited), PR / Media (which press domains cite the brand vs competitors), YouTube (which video creators surface), Reddit (which threads the AI mines).

7. Variance handling & statistical honesty

Single-run metrics are flagged as "low confidence" in the UI. Multi-run metrics surface with their variance band visible (live : the confidence interval is shown as plus-or-minus points on the mention rate, alongside the sample size per metric).

Cross-scan comparisons require a minimum of 7 days between scans to avoid measuring intra-week LLM update noise. Trend metrics use rolling averages with a 30-day window. Model eras (section 3) complete this : change indicators are never computed across a model-version change.

8. Privacy & AI Act posture

sen-ai.fr is classified as limited-risk AI under the EU AI Act (Regulation (EU) 2024/1689), applicable 2 August 2026. We are a downstream deployer of third-party AI systems, not a provider of general-purpose AI. Transparency obligations apply (Article 50) ; no conformity assessment or CE marking.

Personal data exposure is minimal :

  • We store only the email + name of users who access a workspace.
  • AI provider prompts contain brand / domain / topic data only - no end-user PII.
  • All data is hosted in the EU (Hetzner, Helsinki, Finland).
  • US-based AI providers (OpenAI, Anthropic) operate under EU-US Data Privacy Framework + SCC.
  • Organizations can optionally use their own provider API keys (BYOK). In that case, prompts for those providers are sent through the customer's own provider account under the customer's direct agreement with the provider; sen-ai.fr platform keys and the corresponding sub-processor terms apply only when no customer key is configured.

Full disclosure in our Privacy Policy. Logged-in customers can download a per-scan transparency report and an org-level DPIA template from the in-app Compliance section.

8.1 Sub-processors changelog

History of sub-processor additions, sunsets, and scope changes since the initial AI Act disclosure. Updated in lock-step with the in-app compliance pack.

Date Change Sub-processor Notes
2026-07-16 Scope change OpenAI, L.L.C. Option BYOK (cles API fournies par le client) : les organisations peuvent enregistrer leurs propres cles OpenAI, Anthropic, Gemini ou Mistral. Quand une cle client est active, les prompts du fournisseur concerne partent via le compte du client, sous son propre contrat fournisseur. Sans cle client, les cles plateforme sen-ai restent utilisees. Aucun sous-traitant ajoute ni retire - le perimetre de traitement de chaque fournisseur est inchange, seul le compte porteur change.
2026-06-28 Scope change Hetzner Online GmbH Correction de la region d hebergement Hetzner : Helsinki, Finlande (precedemment libelle Falkenstein, Allemagne par erreur). Aucun changement reel : l infrastructure est et reste dans l Union europeenne.
2026-06-28 Added ARCHI301 (RCS Nice 950 897 371) Haloscan ajoute au registre : donnees mots-cles et SERP Google France, amorce des topics et personas. Integration SEO existante, traite des mots-cles et domaines, pas de donnee personnelle.
2026-06-28 Added Babbar Technologies SAS YourTextGuru ajoute au registre : scores semantiques SOSEO et DSEO pour la generation de contenu. Edite par Babbar Technologies SAS, traite des mots-cles de contenu, pas de donnee personnelle.
2026-06-28 Added Apexx LLC (United States) Link Finder ajoute au registre : comparateur de prix du netlinking pour la fonctionnalite alternative media. Editeur Apexx LLC (Etats-Unis), hebergement Contabo (Allemagne), traite des domaines et prix, pas de donnee personnelle.
2026-05-29 Initial disclosure (all) Initial AI Act compliance pack published. Sub-processors registry frozen at 6 entries (Hetzner, OpenAI, Google Ireland, Anthropic, Stripe Europe, Babbar).

9. What sen-ai.fr does NOT do

  • We don't train, fine-tune, or contribute to any AI model.
  • We don't scrape competitor websites at scale beyond what an LLM already cited.
  • We don't impersonate user agents (no Googlebot spoofing, no fake browser strings).
  • We don't use prompt injection or any technique that exploits provider terms of service.
  • We don't engage in any practice prohibited by Article 5 of the EU AI Act (social scoring, real-time biometric ID, exploitative manipulation, etc.).

10. Contact

Compliance / methodology questions : [email protected].

DPO not yet mandated (sen-ai.fr is below the threshold). The above contact handles all data-subject-rights requests and replies within 30 days as required by GDPR.