feat(aep): M2 — worker adapté (Bifrost + scrape léger + ntfy)

Reprend worker/enrich.js (fusion, pas réécriture) pour le nouveau format de
soumission libre et l'infra en prod (décisions MOE du 27/09) :
- LLM Mistral direct → Bifrost (${BIFROST_URL}/v1/chat/completions,
  x-bf-vk), modèle WORKER_MODEL (défaut groq/llama-3.1-8b-instant).
- Scrape crawl4ai/Python → fetch natif Node 22 : timeout 8s, corps
  plafonné 500 Ko, titre + meta description + og:* + texte tronqué 4000c.
- Email Resend → notification ntfy, jamais l'email ni le texte du
  contributeur (id NocoDB, nom suggéré, type, confiance uniquement).
- Sortie LLM {nom, description, type_suggere, ville, tags, confiance} ;
  le worker ne réécrit `nom`/`submission_type` que si le formulaire assoupli
  a laissé un placeholder ("[à qualifier]" / "Type : non précisé").
- Seuil « 5 fiches pending » retiré (décision volume faible, notée si le
  volume remonte).
- --dry-run : fixture locale, mock Bifrost/scrape par défaut, aucune
  écriture NocoDB, DRY_RUN_LIVE=1 pour forcer de vrais appels réseau.
- worker/daily-digest.js supprimé (non référencé, cron purgé le 15/07).

deploy/aep-worker/ : service + timer (15 min) + README avec les étapes
exactes de déploiement pour la session qui déploiera (backup NocoDB,
copie, systemctl, test bout-en-bout M3) — rien exécuté ni copié ici.

PIPE-IA-DOC.md §11 : documente tous les écarts vs la version NAV V2 d'origine.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
Jules
2026-09-27 16:01:53 +02:00
co-authored by Claude Opus 5.5
parent 09d1dc9d5f
commit 695cc8bbc2
7 changed files with 473 additions and 418 deletions
+51
View File
@@ -330,3 +330,54 @@ Pattern à généraliser :
4. Normalize(tags) → taxonomie contrôlée 4. Normalize(tags) → taxonomie contrôlée
5. Log(usage) → stats_usage 5. Log(usage) → stats_usage
``` ```
---
## 11. B5 (2026-09-27) — Formulaire assoupli + pipe adapté (fusion, pas réécriture)
Ce qui suit documente les écarts posés par le batch B5 (décisions de la
maîtrise d'œuvre du 27/09). Les sections 1 à 10 ci-dessus restent la
référence de fond (taxonomie, système prompt d'origine, circuit breaker) ;
cette section documente ce qui a changé dans `worker/enrich.js` et le pipe
de soumission (`server/api/submit/`).
### 11.1 Mapping soumission libre → colonnes NocoDB existantes (M1)
Le formulaire `/proposer` assoupli (`components/FormLibre.vue`,
`utils/submitLibre.ts`) écrit dans la table orgas existante, **aucune colonne
nouvelle** :
| Champ formulaire | Colonne NocoDB | Règle |
|---|---|---|
| 1er lien de la liste | `url` | `null` si aucun lien |
| texte « pourquoi » + tous les liens | `description_user` | `{texte}\n\nLiens :\n{lien1}\n{lien2}...` — et si aucun type suggéré, la ligne `Type : non précisé` est préfixée en tête |
| — | `nom` | hostname du 1er lien, ou 60 premiers caractères du texte si aucun lien, préfixé `[à qualifier] ` |
| chip type suggéré | `submission_type` | valeur choisie (`ecosysteme`/`reseau`/`job`/`outil`) ; si aucune chip → `ecosysteme` par défaut (voir ligne `Type : non précisé` ci-dessus) |
| chip « Références » | *(table `ressources_references`, pas orgas)* | `titre`/`auteur` placeholders, `description` = texte + liens, comme les autres |
| email (optionnel) | `submitted_by_email` | — |
**À faire au déploiement** : ajouter l'option `libre` à la liste de valeurs
acceptées par `submission_type` si l'on veut un jour distinguer une
soumission libre d'une soumission via un des 5 formulaires détaillés — non
fait ici pour ne pas tester un ALTER contre la prod sans l'avoir vérifié.
### 11.2 Worker — écarts (M2)
| Sujet | NAV V2 (§1-10 ci-dessus) | AEP B5 |
|---|---|---|
| LLM | Mistral Nemo, appel direct | **Bifrost** (`${BIFROST_URL}/v1/chat/completions`, header `x-bf-vk`), modèle `WORKER_MODEL` (défaut `groq/llama-3.1-8b-instant`) |
| Scrape | crawl4ai (Python, `AsyncHTTPCrawlerStrategy`) | **fetch natif Node 22** — timeout 8 s, corps plafonné 500 Ko, extraction titre + meta description + `og:*` + texte visible tronqué à 4000 caractères, `User-Agent: AEP/2.0 contact@trans-former.fr` |
| Notification | Resend (email Jules) | **ntfy** (`POST https://ntfy.sh/$NTFY_TOPIC`) — le message ne contient JAMAIS l'email ni le texte libre du contributeur, seulement id NocoDB / nom suggéré / type / confiance |
| Seuil « 5 fiches pending » | Email si ≥ 5 en attente | **Retiré** (décision MOE : volume faible, une notif par fiche traitée suffit — à réactiver si le volume monte) |
| Sortie JSON du LLM | `description_enrichie, points_cles, tags_fonction, echelle, territoire, localisation_ville, confiance` | `{nom, description, type_suggere, ville, tags[], confiance}` — le worker écrit ensuite dans les colonnes existantes (`description_enrichie`, `tags_fonction`, `localisation_ville`) |
| Écriture de `nom` | jamais réécrit | réécrit **seulement** si le nom actuel commence par `[à qualifier]` (placeholder posé par le formulaire assoupli) |
| Écriture de `submission_type` | jamais réécrit | réécrit **seulement** si `description_user` commence par `Type : non précisé` (l'utilisateur n'a pas choisi de chip) et que le type suggéré par le LLM est valide |
| Prix des tokens | fixe (prix Mistral Nemo) | variables `WORKER_PRICE_IN_USD_PER_M` / `WORKER_PRICE_OUT_USD_PER_M`, défaut **0** (tier Groq du tier RAPIDE Bifrost, gratuit) |
| Lock | `/tmp/nav-worker.lock` | `/tmp/aep-worker.lock` (configurable `WORKER_LOCK_FILE`) |
| Chemin VPS | `/opt/nav-carte/worker/` | `/opt/aep-worker/` (voir `deploy/aep-worker/README.md`) |
| Timer | 5 min | 15 min (`deploy/aep-worker/aep-worker.timer`) |
Le mode `--dry-run` (`node enrich.js --dry-run`) lit
`worker/fixtures/dry-run-rows.json` au lieu de NocoDB, n'écrit rien, et par
défaut mock aussi le scrape et l'appel Bifrost (`DRY_RUN_LIVE=1` pour
forcer de vrais appels réseau sans jamais toucher NocoDB).
+90
View File
@@ -0,0 +1,90 @@
# Déploiement aep-worker (B5-M2)
Ce dossier n'a **rien été exécuté ni copié sur le VPS** par la session B5 — code
seulement, comme demandé. Étapes exactes pour la session qui déploiera.
## 0. Pré-requis avant de commencer
- [ ] **Backup NocoDB < 24 h** avant le premier passage du worker (il écrit en base).
- [ ] Vérifier que `/opt/aep/.env` contient déjà : `NOCODB_URL`, `NOCODB_TOKEN`,
`NOCODB_BASE`, `NOCODB_TABLE_ORGAS`, `NOCODB_TABLE_STATS`, `BIFROST_URL`,
`BIFROST_VK`, `NTFY_TOPIC` (et optionnellement `WORKER_MODEL`,
`WORKER_LIMIT`, `BUDGET_MAX_EUR`). Ne PAS créer de nouveau fichier .env —
celui-ci existe déjà en prod (mission dit de le lire, pas de le chercher
dans ce repo).
- [ ] Confirmer que `/opt/nav-carte` (l'ancien worker) est bien absent — il l'était
au 27/09 selon B0. Sinon, l'arrêter/purger avant d'activer celui-ci pour
éviter un double traitement des mêmes fiches pending.
## 1. Copier le code
```bash
ssh vps
mkdir -p /opt/aep-worker
exit
# depuis le poste de dev, après avoir buildé/vérifié la branche :
scp -r worker/enrich.js worker/package.json worker/fixtures vps:/opt/aep-worker/
ssh vps "cd /opt/aep-worker && npm install --omit=dev"
```
(`npm install` : le worker n'a plus de dépendance à `dotenv` côté prod puisque
`EnvironmentFile=` du service systemd charge déjà `/opt/aep/.env` — le
`package.json` du worker peut être simplifié à cette occasion, ou laissé tel
quel si `dotenv` sert encore pour un lancement manuel en local.)
## 2. Poser les unités systemd
```bash
scp deploy/aep-worker/aep-worker.service deploy/aep-worker/aep-worker.timer \
vps:/tmp/
ssh vps
sudo mv /tmp/aep-worker.service /tmp/aep-worker.timer /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now aep-worker.timer
```
## 3. Vérifier
```bash
systemctl status aep-worker.timer
systemctl list-timers | grep aep-worker
# Lancer un passage manuel (sans attendre les 15 min) :
sudo systemctl start aep-worker.service
journalctl -u aep-worker.service -n 100 --no-pager
```
Logs attendus : `=== Worker AEP enrichissement démarré ===`, un compte de
fiches pending, puis soit `Rien à traiter.` soit le détail par fiche
(scrape/Bifrost/ntfy).
## 4. Test bout-en-bout (M3, session suivante)
1. Soumettre une vraie ressource de test via `/proposer` (formulaire assoupli).
2. Vérifier le 201 + le record `pending`/`ai_processed=false` dans NocoDB.
3. Vérifier la notif ntfy de soumission (topic `NTFY_TOPIC`).
4. Lancer le worker (`systemctl start aep-worker.service` ou attendre le
timer) → vérifier `moderation_status=ai_processed`, `description_enrichie`,
`tags_fonction`, `ai_raw_output` remplis, et la notif ntfy « Fiche enrichie ».
5. Modérer dans NocoDB UI (`ai_processed` → `approved`) → vérifier l'affichage
sur le front.
6. **Supprimer les records de test** dans NocoDB (orgas + stats_usage) une
fois la preuve faite.
7. Mettre à jour `PIPE-IA-DOC.md`, `Onglets/Proposer.md`, `PILOTE - AEP.md`
dans le vault avec le résultat réel (pas seulement ce README).
## Anti-chevauchement
Lock fichier PID : `/tmp/aep-worker.lock` (configurable via `WORKER_LOCK_FILE`
dans `.env` si besoin). Nettoyé en fin de run normal, et automatiquement si le
PID contenu dans le lock n'est plus vivant (voir `acquireLock()` dans
`enrich.js`) — pas d'intervention manuelle attendue, sauf lock resté après un
crash dur du process (`kill -9` externe) : dans ce cas `rm /tmp/aep-worker.lock`
suffit.
## Écarts par rapport à la version NAV V2 (PIPE-IA-DOC.md)
Voir l'en-tête de commentaire de `worker/enrich.js` et `PIPE-IA-DOC.md` §11 —
LLM Bifrost au lieu de Mistral direct, scrape fetch natif au lieu de crawl4ai,
notification ntfy au lieu de Resend, écriture `nom`/`submission_type` par le
worker seulement quand le formulaire assoupli a laissé un placeholder.
+11
View File
@@ -0,0 +1,11 @@
[Unit]
Description=AEP — Worker enrichissement IA (scrape léger + Bifrost)
After=network-online.target
Wants=network-online.target
[Service]
Type=oneshot
WorkingDirectory=/opt/aep-worker
EnvironmentFile=/opt/aep/.env
ExecStart=/usr/bin/node enrich.js
TimeoutStartSec=300
+9
View File
@@ -0,0 +1,9 @@
[Unit]
Description=Timer 15 min — worker enrichissement IA AEP
[Timer]
OnCalendar=*:0/15
Persistent=true
[Install]
WantedBy=timers.target
-234
View File
@@ -1,234 +0,0 @@
#!/usr/bin/env node
// worker/daily-digest.js — Email récap quotidien nouvelles fiches AEP
// Cron : 0 8 * * * (tous les jours à 8h)
// Usage : node worker/daily-digest.js
import 'dotenv/config'
// ─── CONFIG DEPUIS .env ───────────────────────────────────────────────────────
const NOCODB_URL = process.env.NOCODB_URL || 'http://localhost:8070';
const NOCODB_TOKEN = process.env.NOCODB_TOKEN;
const NOCODB_BASE = process.env.NOCODB_BASE;
const NOCODB_TABLE_ORGAS = process.env.NOCODB_TABLE_ORGAS;
const RESEND_API_KEY = process.env.RESEND_API_KEY;
const RESEND_FROM = process.env.RESEND_FROM || 'AEP Digest <noreply@trans-former.fr>';
const EMAIL_DEST = process.env.EMAIL_JULES || 'transformationsresilientes@gmail.com';
const NOCODB_ADMIN_URL = process.env.NOCODB_ADMIN_URL || 'http://localhost:8070';
// ─── UTILITAIRES LOG ─────────────────────────────────────────────────────────
function log(...args) {
const ts = new Date().toISOString();
console.log(`[${ts}]`, ...args);
}
// ─── NOCODB — FETCH FICHES DERNIÈRES 24H ─────────────────────────────────────
async function fetchNewFiches() {
const since = new Date(Date.now() - 24 * 60 * 60 * 1000).toISOString();
// NocoDB ne supporte pas bien le filtre datetime sur CreatedAt — on récupère
// les 200 dernières fiches triées par date décroissante et on filtre en JS
// (même pattern que getBudgetMoisCourant dans enrich.js)
const url = `${NOCODB_URL}/api/v1/db/data/noco/${NOCODB_BASE}/${NOCODB_TABLE_ORGAS}?limit=200&sort=-CreatedAt`;
const res = await fetch(url, {
headers: { 'xc-token': NOCODB_TOKEN }
});
if (!res.ok) {
throw new Error(`NocoDB GET fiches → ${res.status}: ${await res.text()}`);
}
const data = await res.json();
const rows = data.list || [];
// Filtre JS sur les 24 dernières heures
const recent = rows.filter(row => {
const created = new Date(row.CreatedAt || row.created_at || 0);
return created >= new Date(since);
});
return recent;
}
// ─── CONSTRUCTION EMAIL HTML ─────────────────────────────────────────────────
function buildEmailHtml(fiches) {
const count = fiches.length;
const date = new Date().toLocaleDateString('fr-FR', {
weekday: 'long', day: 'numeric', month: 'long', year: 'numeric'
});
const lignes = fiches.map(f => {
const nom = f.nom || '(sans nom)';
const url = f.url || null;
const echelle = f.echelle || '—';
const statut = f.moderation_status || 'pending';
const ficheUrl = `https://aep.trans-former.fr/fiche/${f.Id}`;
const nomHtml = url
? `<a href="${url}" style="color:#1a56db;text-decoration:none;">${escHtml(nom)}</a>`
: escHtml(nom);
const statutColor = statut === 'approved' ? '#057a55'
: statut === 'rejected' ? '#c81e1e'
: statut === 'ai_processed' ? '#1a56db'
: '#9ca3af'; // pending / autre
return `
<tr style="border-bottom:1px solid #f3f4f6;">
<td style="padding:10px 12px;font-size:14px;">${nomHtml}</td>
<td style="padding:10px 12px;font-size:13px;color:#6b7280;">${escHtml(echelle)}</td>
<td style="padding:10px 12px;">
<span style="display:inline-block;padding:2px 8px;border-radius:4px;font-size:12px;
font-weight:600;background:#f3f4f6;color:${statutColor};">
${escHtml(statut)}
</span>
</td>
<td style="padding:10px 12px;font-size:12px;">
<a href="${ficheUrl}" style="color:#6b7280;">fiche #${f.Id}</a>
</td>
</tr>`;
}).join('');
return `<!DOCTYPE html>
<html lang="fr">
<head><meta charset="UTF-8"><meta name="viewport" content="width=device-width,initial-scale=1"></head>
<body style="margin:0;padding:0;background:#f9fafb;font-family:-apple-system,BlinkMacSystemFont,'Segoe UI',sans-serif;">
<table width="100%" cellpadding="0" cellspacing="0" style="background:#f9fafb;padding:32px 0;">
<tr><td align="center">
<table width="600" cellpadding="0" cellspacing="0" style="background:#ffffff;border-radius:8px;overflow:hidden;box-shadow:0 1px 3px rgba(0,0,0,0.1);">
<!-- Header -->
<tr>
<td style="background:#1a56db;padding:24px 32px;">
<p style="margin:0;font-size:12px;color:#bfdbfe;text-transform:uppercase;letter-spacing:1px;">AEP — Agences En Pratique</p>
<h1 style="margin:4px 0 0;font-size:20px;color:#ffffff;font-weight:700;">
${count} nouvelle${count > 1 ? 's' : ''} fiche${count > 1 ? 's' : ''} soumise${count > 1 ? 's' : ''}
</h1>
<p style="margin:4px 0 0;font-size:13px;color:#bfdbfe;">${date}</p>
</td>
</tr>
<!-- Corps -->
<tr>
<td style="padding:24px 32px;">
<p style="margin:0 0 16px;font-size:14px;color:#374151;">
${count} fiche${count > 1 ? 's ont été soumises' : ' a été soumise'} dans les dernières 24 heures.
</p>
<!-- Tableau fiches -->
<table width="100%" cellpadding="0" cellspacing="0" style="border-collapse:collapse;border:1px solid #e5e7eb;border-radius:6px;overflow:hidden;">
<thead>
<tr style="background:#f9fafb;">
<th style="padding:10px 12px;text-align:left;font-size:12px;font-weight:600;color:#6b7280;text-transform:uppercase;letter-spacing:0.5px;">Organisation</th>
<th style="padding:10px 12px;text-align:left;font-size:12px;font-weight:600;color:#6b7280;text-transform:uppercase;letter-spacing:0.5px;">Échelle</th>
<th style="padding:10px 12px;text-align:left;font-size:12px;font-weight:600;color:#6b7280;text-transform:uppercase;letter-spacing:0.5px;">Statut</th>
<th style="padding:10px 12px;text-align:left;font-size:12px;font-weight:600;color:#6b7280;text-transform:uppercase;letter-spacing:0.5px;">Lien</th>
</tr>
</thead>
<tbody>
${lignes}
</tbody>
</table>
<!-- CTA modération -->
<div style="margin-top:24px;text-align:center;">
<a href="${NOCODB_ADMIN_URL}"
style="display:inline-block;padding:10px 20px;background:#1a56db;color:#ffffff;
border-radius:6px;font-size:14px;font-weight:600;text-decoration:none;">
Ouvrir NocoDB pour modérer
</a>
</div>
</td>
</tr>
<!-- Footer -->
<tr>
<td style="padding:16px 32px;border-top:1px solid #f3f4f6;">
<p style="margin:0;font-size:11px;color:#9ca3af;text-align:center;">
Digest automatique AEP — envoyé chaque matin à 8h
</p>
</td>
</tr>
</table>
</td></tr>
</table>
</body>
</html>`;
}
// Échappement HTML minimal
function escHtml(str) {
return String(str ?? '')
.replace(/&/g, '&amp;')
.replace(/</g, '&lt;')
.replace(/>/g, '&gt;')
.replace(/"/g, '&quot;');
}
// ─── ENVOI VIA RESEND ─────────────────────────────────────────────────────────
async function sendDigest(fiches) {
const count = fiches.length;
const subject = `AEP — ${count} nouvelle${count > 1 ? 's' : ''} fiche${count > 1 ? 's' : ''} soumise${count > 1 ? 's' : ''}`;
const html = buildEmailHtml(fiches);
const res = await fetch('https://api.resend.com/emails', {
method: 'POST',
headers: {
'Authorization': `Bearer ${RESEND_API_KEY}`,
'Content-Type': 'application/json'
},
body: JSON.stringify({
from: RESEND_FROM,
to: [EMAIL_DEST],
subject,
html
})
});
if (!res.ok) {
const err = await res.text();
throw new Error(`Resend API ${res.status}: ${err}`);
}
const data = await res.json();
log(`Email digest envoyé → ${EMAIL_DEST} | id: ${data.id} | sujet: "${subject}"`);
}
// ─── MAIN ─────────────────────────────────────────────────────────────────────
async function run() {
log('=== daily-digest démarré ===');
// Vérification variables obligatoires
if (!NOCODB_TOKEN || !NOCODB_BASE || !NOCODB_TABLE_ORGAS) {
log('ERREUR : Variables NocoDB manquantes (NOCODB_TOKEN, NOCODB_BASE, NOCODB_TABLE_ORGAS)');
process.exit(1);
}
if (!RESEND_API_KEY) {
log('ERREUR : RESEND_API_KEY manquante');
process.exit(1);
}
try {
// Fetch fiches des 24 dernières heures
const fiches = await fetchNewFiches();
log(`${fiches.length} fiche(s) trouvée(s) dans les dernières 24h`);
if (fiches.length === 0) {
log('Aucune nouvelle fiche — digest non envoyé.');
process.exit(0);
}
// Envoi email
await sendDigest(fiches);
log('=== daily-digest terminé avec succès ===');
} catch (e) {
log('ERREUR daily-digest:', e.message);
console.error(e.stack);
process.exit(1);
}
}
run();
+286 -184
View File
@@ -1,14 +1,28 @@
#!/usr/bin/env node #!/usr/bin/env node
/** /**
* NAV V2 — Worker enrichissement IA * AEP — Worker enrichissement IA (B5-M2)
* Lancé via systemd timer toutes les 5 minutes * Lancé via systemd timer toutes les 15 minutes (deploy/aep-worker/aep-worker.timer)
* Pipeline : fetch pending → scrape crawl4ai → Mistral Nemo → update NocoDB → log stats * Pipeline : fetch pending → scrape léger (fetch natif) → Bifrost (LLM) → update NocoDB → ntfy
*
* Écarts à la version NAV V2 d'origine (PIPE-IA-DOC.md, fusionné pas réécrit) :
* - LLM : Bifrost (passerelle OpenAI-compatible en prod) au lieu de Mistral direct.
* - Scrape : fetch natif Node 22 (timeout 8s, 500 Ko max, texte tronqué 4000 car.)
* au lieu de crawl4ai/Python (disque VPS tendu, décision MOE 27/09).
* - Notification : ntfy (https://ntfy.sh/$NTFY_TOPIC) au lieu de Resend (abandonné 15/07).
* Le message ne contient JAMAIS l'email ni le texte libre du contributeur.
* - Le formulaire assoupli (B5-M1) écrit un `nom` placeholder ("[à qualifier] host") et
* éventuellement "Type : non précisé" en tête de description_user. Le worker propose un
* vrai nom et un type — voir §Sortie JSON ci-dessous et PIPE-IA-DOC.md.
* - Seuil « email à 5 fiches pending » PAS réactivé (décision MOE : volume faible, une
* notif par fiche traitée suffit — cf. B5-proposer-pipe.md §Points d'attention).
*/ */
import { spawnSync, execSync } from 'child_process'; import { spawnSync, execSync } from 'child_process';
import { existsSync, writeFileSync, unlinkSync } from 'fs'; import { existsSync, writeFileSync, unlinkSync, readFileSync, mkdirSync } from 'fs';
import { tmpdir } from 'os'; import { dirname, join } from 'path';
import { join } from 'path'; import { fileURLToPath } from 'url';
const __dirname = dirname(fileURLToPath(import.meta.url));
// ─── CONFIG DEPUIS .env ─────────────────────────────────────────────────────── // ─── CONFIG DEPUIS .env ───────────────────────────────────────────────────────
const NOCODB_URL = process.env.NOCODB_URL || 'http://localhost:8070'; const NOCODB_URL = process.env.NOCODB_URL || 'http://localhost:8070';
@@ -16,25 +30,44 @@ const NOCODB_TOKEN = process.env.NOCODB_TOKEN;
const NOCODB_BASE = process.env.NOCODB_BASE; const NOCODB_BASE = process.env.NOCODB_BASE;
const NOCODB_TABLE_ORGAS = process.env.NOCODB_TABLE_ORGAS; const NOCODB_TABLE_ORGAS = process.env.NOCODB_TABLE_ORGAS;
const NOCODB_TABLE_STATS = process.env.NOCODB_TABLE_STATS; const NOCODB_TABLE_STATS = process.env.NOCODB_TABLE_STATS;
const MISTRAL_API_KEY = process.env.MISTRAL_API_KEY;
const RESEND_API_KEY = process.env.RESEND_API_KEY; const BIFROST_URL = process.env.BIFROST_URL || 'http://127.0.0.1:8080';
const RESEND_FROM = process.env.RESEND_FROM || 'contact@trans-former.fr'; const BIFROST_VK = process.env.BIFROST_VK;
const EMAIL_JULES = process.env.EMAIL_JULES || 'jules@trans-former.fr'; const WORKER_MODEL = process.env.WORKER_MODEL || 'groq/llama-3.1-8b-instant';
const NTFY_TOPIC = process.env.NTFY_TOPIC;
const BUDGET_MAX_EUR = parseFloat(process.env.BUDGET_MAX_EUR || '20'); const BUDGET_MAX_EUR = parseFloat(process.env.BUDGET_MAX_EUR || '20');
const WORKER_LIMIT = parseInt(process.env.WORKER_LIMIT || '5'); const WORKER_LIMIT = parseInt(process.env.WORKER_LIMIT || '5');
const LOCK_FILE = '/tmp/nav-worker.lock'; const LOCK_FILE = process.env.WORKER_LOCK_FILE || '/tmp/aep-worker.lock';
// ─── PRIX MISTRAL NEMO (USD → EUR) ─────────────────────────────────────────── // Prix par 1M tokens (USD) — par défaut 0 (Groq gratuit dans le tier RAPIDE Bifrost).
const NEMO_PRICE_IN = 0.02 / 1_000_000; // $0.02 / 1M tokens input // Réglable si le worker bascule un jour sur un modèle payant.
const NEMO_PRICE_OUT = 0.04 / 1_000_000; // $0.04 / 1M tokens output const WORKER_PRICE_IN = parseFloat(process.env.WORKER_PRICE_IN_USD_PER_M || '0') / 1_000_000;
const WORKER_PRICE_OUT = parseFloat(process.env.WORKER_PRICE_OUT_USD_PER_M || '0') / 1_000_000;
const USD_TO_EUR = 0.93; const USD_TO_EUR = 0.93;
// Scrape léger
const SCRAPE_TIMEOUT_MS = 8_000;
const SCRAPE_MAX_BYTES = 500_000;
const SCRAPE_TEXT_MAX_CHARS = 4_000;
const SCRAPE_USER_AGENT = 'AEP/2.0 contact@trans-former.fr';
// ─── CLI ──────────────────────────────────────────────────────────────────────
const DRY_RUN = process.argv.includes('--dry-run');
// DRY_RUN_LIVE=1 : en dry-run, tente quand même de vrais appels réseau (scrape + Bifrost)
// pour un test manuel avec accès réseau — ne touche JAMAIS NocoDB en dry-run, dans tous les cas.
const DRY_RUN_LIVE = process.env.DRY_RUN_LIVE === '1';
const DRY_RUN_FIXTURE = process.env.DRY_RUN_FIXTURE || join(__dirname, 'fixtures', 'dry-run-rows.json');
// ─── TAXONOMIE VALIDE (apostrophe typographique U+2019 comme NocoDB) ───────── // ─── TAXONOMIE VALIDE (apostrophe typographique U+2019 comme NocoDB) ─────────
const VALID_FONCTIONS = [ const VALID_FONCTIONS = [
'Juridique', 'Technique', 'Économique', 'Administratif', 'Chantier', 'Juridique', 'Technique', 'Économique', 'Administratif', 'Chantier',
'Comptabilité', 'Développement', 'Formation', 'Gestion d\u2019agence', 'Santé mentale' 'Comptabilité', 'Développement', 'Formation', 'Gestion d’agence', 'Santé mentale'
]; ];
const VALID_SUBMISSION_TYPES = ['ecosysteme', 'reseau', 'job', 'outil'];
// ─── MAPPING NORMALISATION TAGS ─────────────────────────────────────────────── // ─── MAPPING NORMALISATION TAGS ───────────────────────────────────────────────
const TAG_MAP = [ const TAG_MAP = [
[['juridique', 'droit', 'litige', 'contrat', 'déontologie', 'décennale', 'médiation', 'pi ', 'propriété intellectuelle', 'ccag', 'marchés publics droit'], 'Juridique'], [['juridique', 'droit', 'litige', 'contrat', 'déontologie', 'décennale', 'médiation', 'pi ', 'propriété intellectuelle', 'ccag', 'marchés publics droit'], 'Juridique'],
@@ -45,16 +78,14 @@ const TAG_MAP = [
[['comptabilité', 'fiscal', 'tva', 'bnc', 'bic', 'expert-comptable', 'bilan', 'trésorerie', 'transmission agence', 'création agence', 'micro'], 'Comptabilité'], [['comptabilité', 'fiscal', 'tva', 'bnc', 'bic', 'expert-comptable', 'bilan', 'trésorerie', 'transmission agence', 'création agence', 'micro'], 'Comptabilité'],
[['développement', 'prospection', 'commercial', 'client', 'réseau', 'candidature', 'consultation', 'acquisition', 'marketing', 'notoriété', 'ao '], 'Développement'], [['développement', 'prospection', 'commercial', 'client', 'réseau', 'candidature', 'consultation', 'acquisition', 'marketing', 'notoriété', 'ao '], 'Développement'],
[['formation', 'école', 'mooc', 'organisme', 'formation continue', 'cpf', 'dpc', 'cfaa'], 'Formation'], [['formation', 'école', 'mooc', 'organisme', 'formation continue', 'cpf', 'dpc', 'cfaa'], 'Formation'],
[['gestion d\u2019agence', 'gestion d\'agence', 'rh', 'recrutement', 'emploi', 'salaire', 'ccn', 'convention collective', 'idcc', 'temps de travail', 'management'], 'Gestion d\u2019agence'], [['gestion d’agence', 'gestion d\'agence', 'rh', 'recrutement', 'emploi', 'salaire', 'ccn', 'convention collective', 'idcc', 'temps de travail', 'management'], 'Gestion d’agence'],
[['santé mentale', 'burn-out', 'épuisement', 'souffrance', 'bien-être', 'harcèlement', 'stress', 'psychologique', 'équilibre'], 'Santé mentale'], [['santé mentale', 'burn-out', 'épuisement', 'souffrance', 'bien-être', 'harcèlement', 'stress', 'psychologique', 'équilibre'], 'Santé mentale'],
]; ];
function normalizeTag(raw) { function normalizeTag(raw) {
const t = raw.toLowerCase().trim(); const t = String(raw).toLowerCase().trim();
// D'abord chercher correspondance exacte dans les valeurs valides
const exact = VALID_FONCTIONS.find(v => v.toLowerCase() === t); const exact = VALID_FONCTIONS.find(v => v.toLowerCase() === t);
if (exact) return exact; if (exact) return exact;
// Sinon chercher par mots-clés
for (const [patterns, normalized] of TAG_MAP) { for (const [patterns, normalized] of TAG_MAP) {
if (patterns.some(p => t.includes(p))) return normalized; if (patterns.some(p => t.includes(p))) return normalized;
} }
@@ -117,8 +148,18 @@ async function nocodbPost(tableId, data) {
return res.json(); return res.json();
} }
/** Écrit une mise à jour de fiche — no-op loggé en dry-run (jamais de PATCH réel). */
async function patchRow(rowId, data) {
if (DRY_RUN) {
log(`[dry-run] PATCH fiche ${rowId} :`, JSON.stringify(data));
return;
}
return nocodbPatch(NOCODB_TABLE_ORGAS, rowId, data);
}
// ─── BUDGET CIRCUIT BREAKER ────────────────────────────────────────────────── // ─── BUDGET CIRCUIT BREAKER ──────────────────────────────────────────────────
async function getBudgetMoisCourant() { async function getBudgetMoisCourant() {
if (DRY_RUN) return 0;
const now = new Date(); const now = new Date();
const year = now.getFullYear(); const year = now.getFullYear();
const month = now.getMonth(); // 0-indexed const month = now.getMonth(); // 0-indexed
@@ -142,7 +183,12 @@ async function getBudgetMoisCourant() {
async function logUsage(usage, model, endpoint, orgaId) { async function logUsage(usage, model, endpoint, orgaId) {
const tokensIn = usage?.prompt_tokens || 0; const tokensIn = usage?.prompt_tokens || 0;
const tokensOut = usage?.completion_tokens || 0; const tokensOut = usage?.completion_tokens || 0;
const coutEur = ((tokensIn * NEMO_PRICE_IN) + (tokensOut * NEMO_PRICE_OUT)) * USD_TO_EUR; const coutEur = ((tokensIn * WORKER_PRICE_IN) + (tokensOut * WORKER_PRICE_OUT)) * USD_TO_EUR;
if (DRY_RUN) {
log(`[dry-run] Usage (non loggé) : ${tokensIn}in + ${tokensOut}out = €${coutEur.toFixed(6)} (${model})`);
return coutEur;
}
await nocodbPost(NOCODB_TABLE_STATS, { await nocodbPost(NOCODB_TABLE_STATS, {
model, model,
@@ -166,111 +212,168 @@ async function fetchPendingRows() {
return data.list || []; return data.list || [];
} }
// ─── SCRAPING CRAWL4AI (mode HTTP statique, sans Playwright) ───────────────── function loadFixtureRows() {
async function scrapeWithCrawl4ai(url) { if (!existsSync(DRY_RUN_FIXTURE)) {
log(`Scraping: ${url}`); throw new Error(`Fixture dry-run introuvable : ${DRY_RUN_FIXTURE}`);
}
const raw = JSON.parse(readFileSync(DRY_RUN_FIXTURE, 'utf-8'));
return Array.isArray(raw) ? raw : [raw];
}
// Script Python temporaire pour crawl4ai // ─── LIENS — extraits de url + description_user (format B5-M1 "Liens :\n...") ─
const urlSafe = url.replace(/\\/g, '\\\\').replace(/'/g, "\\'"); function extractLinks(row) {
const script = ` const text = `${row.url || ''}\n${row.description_user || row.description || ''}`;
import asyncio, sys const matches = text.match(/https?:\/\/[^\s)"'<>]+/g) || [];
from crawl4ai import AsyncWebCrawler, CrawlerRunConfig const cleaned = matches.map(u => u.replace(/[.,;:!?]+$/, ''));
from crawl4ai.async_crawler_strategy import AsyncHTTPCrawlerStrategy return [...new Set(cleaned)];
}
async def scrape():
strategy = AsyncHTTPCrawlerStrategy()
run_cfg = CrawlerRunConfig(
word_count_threshold=20,
excluded_tags=['nav', 'footer', 'script', 'style', 'head'],
remove_overlay_elements=True
)
async with AsyncWebCrawler(crawler_strategy=strategy, verbose=False) as crawler:
result = await crawler.arun(url='${urlSafe}', config=run_cfg)
if result.success and result.markdown:
content = result.markdown[:16000]
sys.stdout.buffer.write(content.encode('utf-8'))
else:
sys.stderr.write(f"Scrape failed: success={result.success}\\n")
sys.exit(1)
asyncio.run(scrape())
`;
const scriptPath = join(tmpdir(), `nav-scrape-${Date.now()}.py`);
writeFileSync(scriptPath, script, 'utf-8');
// ─── SCRAPE LÉGER (fetch natif, pas de crawl4ai) ─────────────────────────────
async function fetchCapped(url) {
const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), SCRAPE_TIMEOUT_MS);
try { try {
const result = spawnSync('python3', [scriptPath], { const res = await fetch(url, {
timeout: 180_000, signal: controller.signal,
maxBuffer: 20 * 1024 * 1024, redirect: 'follow',
env: { ...process.env, PYTHONDONTWRITEBYTECODE: '1', PYTHONIOENCODING: 'utf-8' } headers: { 'User-Agent': SCRAPE_USER_AGENT },
}); });
if (!res.ok) throw new Error(`HTTP ${res.status}`);
if (result.error) throw result.error; const reader = res.body?.getReader();
if (!reader) return await res.text();
const stdout = result.stdout?.toString('utf-8') || ''; const chunks = [];
const stderr = result.stderr?.toString('utf-8') || ''; let total = 0;
while (true) {
if (result.status === 0 && stdout.length > 50) { const { done, value } = await reader.read();
log(`Scrape OK: ${stdout.length} chars`); if (done) break;
return stdout.trim(); chunks.push(value);
} else { total += value.length;
throw new Error(`Scrape failed (code ${result.status}): ${stderr.slice(0, 300)}`); if (total >= SCRAPE_MAX_BYTES) {
await reader.cancel().catch(() => {});
break;
} }
}
const buf = Buffer.concat(chunks.map(c => Buffer.from(c)));
return buf.subarray(0, SCRAPE_MAX_BYTES).toString('utf-8');
} finally { } finally {
try { unlinkSync(scriptPath); } catch {} clearTimeout(timer);
} }
} }
// ─── APPEL MISTRAL NEMO ─────────────────────────────────────────────────────── function extractMeta(html, key) {
const SYSTEM_PROMPT = `Tu es un assistant spécialisé dans l'écosystème professionnel de l'architecture en France. Tu reçois des informations sur une organisation ou ressource liée au secteur de l'architecture, et tu dois les enrichir pour alimenter une cartographie collaborative. const re1 = new RegExp(`<meta[^>]*(?:name|property)=["']${key}["'][^>]*content=["']([^"']*)["']`, 'i');
const re2 = new RegExp(`<meta[^>]*content=["']([^"']*)["'][^>]*(?:name|property)=["']${key}["']`, 'i');
return (html.match(re1) || html.match(re2))?.[1]?.trim() || null;
}
function extractOgTags(html) {
const og = {};
const re1 = /<meta[^>]*property=["']og:([a-zA-Z:_-]+)["'][^>]*content=["']([^"']*)["']/gi;
const re2 = /<meta[^>]*content=["']([^"']*)["'][^>]*property=["']og:([a-zA-Z:_-]+)["']/gi;
let m;
while ((m = re1.exec(html))) og[m[1]] = m[2];
while ((m = re2.exec(html))) if (!(m[2] in og)) og[m[2]] = m[1];
return og;
}
function decodeEntities(s) {
return s
.replace(/&nbsp;/g, ' ')
.replace(/&amp;/g, '&')
.replace(/&lt;/g, '<')
.replace(/&gt;/g, '>')
.replace(/&quot;/g, '"')
.replace(/&#0?39;/g, "'");
}
function extractVisibleText(html) {
let s = html
.replace(/<!--[\s\S]*?-->/g, ' ')
.replace(/<script[\s\S]*?<\/script>/gi, ' ')
.replace(/<style[\s\S]*?<\/style>/gi, ' ')
.replace(/<nav[\s\S]*?<\/nav>/gi, ' ')
.replace(/<footer[\s\S]*?<\/footer>/gi, ' ')
.replace(/<head[\s\S]*?<\/head>/gi, ' ');
s = s.replace(/<[^>]+>/g, ' ');
s = decodeEntities(s);
s = s.replace(/\s+/g, ' ').trim();
return s.slice(0, SCRAPE_TEXT_MAX_CHARS);
}
function extractTitle(html) {
return html.match(/<title[^>]*>([^<]*)<\/title>/i)?.[1]?.trim() || null;
}
async function scrapeLight(url) {
log(`Scraping léger: ${url}`);
const html = await fetchCapped(url);
return {
title: extractTitle(html),
metaDescription: extractMeta(html, 'description'),
og: extractOgTags(html),
text: extractVisibleText(html),
};
}
const MOCK_SCRAPE_RESULT = {
title: '[mock dry-run] Titre de la page',
metaDescription: '[mock dry-run] Meta description factice, aucun réseau contacté.',
og: { site_name: '[mock dry-run]' },
text: '[mock dry-run] Extrait de texte factice utilisé quand DRY_RUN_LIVE n’est pas activé.',
};
// ─── APPEL LLM VIA BIFROST ────────────────────────────────────────────────────
const SYSTEM_PROMPT = `Tu es un assistant qui aide à qualifier des ressources soumises pour une cartographie collaborative de l'écosystème professionnel de l'architecture en France (AEP).
RÈGLES ABSOLUES : RÈGLES ABSOLUES :
1. Tu ne dois JAMAIS inventer d'informations non présentes dans les sources fournies. 1. Tu ne dois JAMAIS inventer d'informations non présentes dans les sources fournies.
2. Si une information est absente ou incertaine, retourne \`null\` pour ce champ. 2. Si une information est absente ou incertaine, retourne \`null\` pour ce champ.
3. Tu dois retourner UNIQUEMENT un objet JSON valide, sans texte avant ou après. 3. Tu dois retourner UNIQUEMENT un objet JSON valide, sans texte avant ou après.
4. La description_enrichie doit être neutre, factuelle, en français, sans jugement de valeur. 4. "description" : neutre, factuelle, en français, max 300 caractères, sans jugement de valeur.
5. Les points_cles sont des phrases courtes (max 12 mots chacune), actionnables pour un architecte. 5. "nom" : un nom de fiche court et identifiable, sans le préfixe "[à qualifier]".
6. Pour les tags_fonction, ne propose que des valeurs parmi la liste autorisée. 6. "tags" : 1 à 5 valeurs, uniquement parmi la liste autorisée.
TAXONOMIE AUTORISÉE : TAGS AUTORISÉS : "Juridique" | "Technique" | "Économique" | "Administratif" | "Chantier" | "Comptabilité" | "Développement" | "Formation" | "Gestion d’agence" | "Santé mentale"
- Échelle (une seule valeur) : "National" | "Régional" | "Départemental" | "Local" TYPE_SUGGERE AUTORISÉ (une seule valeur ou null) : "ecosysteme" | "reseau" | "job" | "outil"
- Territoire (une seule valeur) : "Métropole" | "Guadeloupe" | "Martinique" | "Guyane" | "Réunion" | "Mayotte" | null
- Tags fonction (1 à 5 valeurs) : "Juridique" | "Technique" | "Économique" | "Administratif" | "Chantier" | "Comptabilité" | "Développement" | "Formation" | "Gestion d\u2019agence" | "Santé mentale"
FORMAT DE SORTIE JSON : FORMAT DE SORTIE JSON :
{ {
"description_enrichie": "string (max 300 chars, français, neutre, factuel)", "nom": "string | null",
"points_cles": ["string", "string", "string"], "description": "string (max 300 chars, français, neutre, factuel) | null",
"tags_fonction": ["Valeur1", "Valeur2"], "type_suggere": "ecosysteme" | "reseau" | "job" | "outil" | null,
"echelle": "National" | "Régional" | "Départemental" | "Local" | null, "ville": "string | null",
"territoire": "Métropole" | ... | null, "tags": ["string", "..."],
"localisation_ville": "string" | null,
"confiance": "haute" | "moyenne" | "faible" "confiance": "haute" | "moyenne" | "faible"
} }
Le champ "confiance" reflète ta certitude globale sur l'enrichissement : Le champ "confiance" reflète ta certitude globale :
- "haute" : URL scrapée avec contenu riche, informations claires - "haute" : contenu scrapé riche, informations claires.
- "moyenne" : URL scrapée mais contenu partiel, ou description_user seule suffisante - "moyenne" : contenu partiel, ou texte du contributeur seul mais clair.
- "faible" : URL non disponible et description_user vague, inférences importantes`; - "faible" : aucun contenu scrapé et texte du contributeur vague, inférences importantes.`;
function buildUserPrompt(row, scrapeContent) { function buildUserPrompt(row, scrapeData, autresLiens) {
return `ORGANISATION À ENRICHIR : return `RESSOURCE À QUALIFIER :
Nom : ${row.nom} Nom actuel (placeholder à remplacer) : ${row.nom}
URL : ${row.url || 'non fournie'} Lien principal : ${row.url || 'non fourni'}
Description soumise par l'utilisateur : ${row.description_user || row.description || 'non fournie'} ${autresLiens.length ? `Autres liens mentionnés par le contributeur : ${autresLiens.join(', ')}` : ''}
Texte du contributeur (pourquoi c'est pertinent pour lui) : ${row.description_user || row.description || 'non fourni'}
CONTENU DU SITE WEB (extrait par scraping) : CONTENU EXTRAIT DU SITE (scraping léger) :
${scrapeContent || 'Site non accessible ou URL non fournie.'} Titre : ${scrapeData?.title || 'non disponible'}
Meta description : ${scrapeData?.metaDescription || 'non disponible'}
Open Graph : ${scrapeData && Object.keys(scrapeData.og || {}).length ? JSON.stringify(scrapeData.og) : 'non disponible'}
Texte visible (extrait) : ${scrapeData?.text || 'Site non accessible ou lien non fourni.'}
--- ---
Enrichis cette fiche selon les règles du system prompt. Retourne uniquement le JSON.`; Qualifie cette ressource selon les règles du system prompt. Retourne uniquement le JSON.`;
} }
async function callMistralWithRetry(row, scrapeContent, maxRetries = 2) { async function callBifrostWithRetry(row, scrapeData, autresLiens, maxRetries = 2) {
const userPrompt = buildUserPrompt(row, scrapeContent); const userPrompt = buildUserPrompt(row, scrapeData, autresLiens);
for (let attempt = 0; attempt <= maxRetries; attempt++) { for (let attempt = 0; attempt <= maxRetries; attempt++) {
if (attempt > 0) { if (attempt > 0) {
@@ -279,14 +382,14 @@ async function callMistralWithRetry(row, scrapeContent, maxRetries = 2) {
} }
try { try {
const res = await fetch('https://api.mistral.ai/v1/chat/completions', { const res = await fetch(`${BIFROST_URL}/v1/chat/completions`, {
method: 'POST', method: 'POST',
headers: { headers: {
'Authorization': `Bearer ${MISTRAL_API_KEY}`, 'x-bf-vk': BIFROST_VK,
'Content-Type': 'application/json' 'Content-Type': 'application/json'
}, },
body: JSON.stringify({ body: JSON.stringify({
model: 'open-mistral-nemo', model: WORKER_MODEL,
temperature: 0.2, temperature: 0.2,
max_tokens: 800, max_tokens: 800,
response_format: { type: 'json_object' }, response_format: { type: 'json_object' },
@@ -300,101 +403,88 @@ async function callMistralWithRetry(row, scrapeContent, maxRetries = 2) {
if (!res.ok) { if (!res.ok) {
const err = await res.text(); const err = await res.text();
throw new Error(`Mistral API ${res.status}: ${err}`); throw new Error(`Bifrost API ${res.status}: ${err}`);
} }
const data = await res.json(); const data = await res.json();
const content = data.choices?.[0]?.message?.content; const content = data.choices?.[0]?.message?.content;
if (!content) throw new Error('Réponse Mistral vide'); if (!content) throw new Error('Réponse Bifrost vide');
const parsed = JSON.parse(content); const parsed = JSON.parse(content);
// Attacher usage pour logging
parsed._usage = data.usage; parsed._usage = data.usage;
parsed._raw = content; parsed._raw = content;
return parsed; return parsed;
} catch (e) { } catch (e) {
log(`Erreur Mistral (tentative ${attempt + 1}): ${e.message}`); log(`Erreur Bifrost (tentative ${attempt + 1}): ${e.message}`);
if (attempt === maxRetries) return null; if (attempt === maxRetries) return null;
} }
} }
return null; return null;
} }
// ─── EMAIL JULES VIA RESEND ─────────────────────────────────────────────────── const MOCK_BIFROST_RESULT = {
async function sendEmailJules(subject, body) { nom: '[mock dry-run] Nom suggéré',
if (!RESEND_API_KEY) { description: '[mock dry-run] Description factice générée sans appel réseau.',
log('RESEND_API_KEY absent, email skippé'); type_suggere: 'ecosysteme',
ville: null,
tags: ['Développement'],
confiance: 'faible',
_usage: { prompt_tokens: 0, completion_tokens: 0 },
_raw: '{"mock":"dry-run"}',
};
// ─── NOTIFICATION NTFY (jamais l'email ni le texte du contributeur) ──────────
async function notifyNtfy(title, message) {
if (DRY_RUN) {
log(`[dry-run] ntfy (non envoyé) — ${title} :`, message.replace(/\n/g, ' | '));
return;
}
if (!NTFY_TOPIC) {
log('NTFY_TOPIC absent, notification skippée');
return; return;
} }
try { try {
const res = await fetch('https://api.resend.com/emails', { const res = await fetch(`https://ntfy.sh/${NTFY_TOPIC}`, {
method: 'POST', method: 'POST',
headers: { headers: { Title: title, Priority: '3', Tags: 'inbox_tray' },
'Authorization': `Bearer ${RESEND_API_KEY}`, body: message,
'Content-Type': 'application/json'
},
body: JSON.stringify({
from: RESEND_FROM,
to: [EMAIL_JULES],
subject,
text: body
})
}); });
if (!res.ok) log('Email error:', await res.text()); if (!res.ok) log('ntfy erreur:', await res.text());
else log('Email envoyé à Jules:', subject); else log('ntfy envoyé:', title);
} catch (e) { } catch (e) {
log('Email exception:', e.message); log('ntfy exception:', e.message);
}
}
// ─── VÉRIFICATION SEUIL 5 FICHES PENDING MODÉRATION ─────────────────────────
async function checkModerationQueue() {
const data = await nocodbGet(
`${NOCODB_TABLE_ORGAS}?where=(moderation_status,eq,ai_processed)&limit=100`
);
const count = data.pageInfo?.totalRows || 0;
if (count >= 5) {
log(`Seuil modération atteint: ${count} fiches en attente`);
await sendEmailJules(
`NAV — ${count} fiches à modérer`,
`Bonjour Jules,\n\n${count} fiches ont été enrichies par l'IA et attendent ta validation dans NocoDB.\n\nLien NocoDB : http://localhost:8070\nFiltre : moderation_status = ai_processed\n\nBonne modération !`
);
} }
} }
// ─── MAIN ───────────────────────────────────────────────────────────────────── // ─── MAIN ─────────────────────────────────────────────────────────────────────
async function run() { async function run() {
if (!acquireLock()) { if (!DRY_RUN && !acquireLock()) {
log('Worker déjà en cours, skip'); log('Worker déjà en cours, skip');
process.exit(0); process.exit(0);
} }
const startTime = Date.now(); const startTime = Date.now();
log('=== Worker NAV enrichissement démarré ==='); log(`=== Worker AEP enrichissement démarré${DRY_RUN ? ' (--dry-run)' : ''} ===`);
try { try {
// Vérification clés obligatoires if (!DRY_RUN) {
if (!NOCODB_TOKEN || !NOCODB_BASE || !NOCODB_TABLE_ORGAS || !MISTRAL_API_KEY) { if (!NOCODB_TOKEN || !NOCODB_BASE || !NOCODB_TABLE_ORGAS || !BIFROST_VK) {
throw new Error('Variables .env manquantes (NOCODB_TOKEN, NOCODB_BASE, NOCODB_TABLE_ORGAS, MISTRAL_API_KEY)'); throw new Error('Variables .env manquantes (NOCODB_TOKEN, NOCODB_BASE, NOCODB_TABLE_ORGAS, BIFROST_VK)');
}
} }
// Check budget global
const budgetMois = await getBudgetMoisCourant(); const budgetMois = await getBudgetMoisCourant();
log(`Budget mois courant: €${budgetMois.toFixed(4)} / €${BUDGET_MAX_EUR}`); log(`Budget mois courant: €${budgetMois.toFixed(4)} / €${BUDGET_MAX_EUR}`);
if (budgetMois >= BUDGET_MAX_EUR) { if (budgetMois >= BUDGET_MAX_EUR) {
log('Budget épuisé pour ce mois. Worker en pause.'); log('Budget épuisé pour ce mois. Worker en pause.');
await sendEmailJules( await notifyNtfy('[AEP] Budget IA épuisé', `Le budget de ${BUDGET_MAX_EUR}€ a été atteint. Worker en pause jusqu'au 1er du mois prochain.`);
'NAV — Budget IA épuisé ce mois',
`Le budget IA de ${BUDGET_MAX_EUR}€ a été atteint. Le worker est en pause jusqu'au 1er du mois prochain.\n\nConsommation actuelle : €${budgetMois.toFixed(4)}`
);
return; return;
} }
// Fetch fiches pending const rows = DRY_RUN ? loadFixtureRows() : await fetchPendingRows();
const rows = await fetchPendingRows(); log(`${rows.length} fiche(s) à traiter${DRY_RUN ? ' (fixture)' : ''}`);
log(`${rows.length} fiche(s) à traiter`);
if (rows.length === 0) { if (rows.length === 0) {
log('Rien à traiter.'); log('Rien à traiter.');
@@ -407,71 +497,86 @@ async function run() {
const rowStart = Date.now(); const rowStart = Date.now();
log(`--- Traitement fiche ${row.Id}: ${row.nom} ---`); log(`--- Traitement fiche ${row.Id}: ${row.nom} ---`);
// Re-check budget avant chaque fiche
const budgetCheck = await getBudgetMoisCourant(); const budgetCheck = await getBudgetMoisCourant();
if (budgetCheck >= BUDGET_MAX_EUR) { if (budgetCheck >= BUDGET_MAX_EUR) {
log('Budget atteint mid-pipeline, arrêt.'); log('Budget atteint mid-pipeline, arrêt.');
break; break;
} }
// Scraping const liens = extractLinks(row);
let scrapeContent = null; const primaryUrl = row.url && row.url.trim() ? row.url.trim() : null;
const hasUrl = row.url && row.url.trim().length > 0; const autresLiens = liens.filter(l => l !== primaryUrl);
const shouldScrape = hasUrl && (row.scrape_status === 'pending' || !row.scrape_status);
// Scraping (uniquement le lien principal — un seul champ scrape_content en base)
let scrapeData = null;
const shouldScrape = primaryUrl && (row.scrape_status === 'pending' || !row.scrape_status);
const useNetwork = !DRY_RUN || DRY_RUN_LIVE;
if (shouldScrape) { if (shouldScrape) {
try { try {
scrapeContent = await scrapeWithCrawl4ai(row.url); scrapeData = useNetwork ? await scrapeLight(primaryUrl) : MOCK_SCRAPE_RESULT;
await nocodbPatch(NOCODB_TABLE_ORGAS, row.Id, { await patchRow(row.Id, {
scrape_status: 'scraped', scrape_status: 'scraped',
scrape_content: scrapeContent scrape_content: JSON.stringify(scrapeData),
}); });
} catch (e) { } catch (e) {
log(`Scrape échoué: ${e.message}`); log(`Scrape échoué: ${e.message}`);
await nocodbPatch(NOCODB_TABLE_ORGAS, row.Id, { scrape_status: 'failed' }); await patchRow(row.Id, { scrape_status: 'failed' });
// Continue avec l'IA sans contenu scrape
} }
} else if (!hasUrl) { } else if (!primaryUrl) {
await nocodbPatch(NOCODB_TABLE_ORGAS, row.Id, { scrape_status: 'no_link' }); await patchRow(row.Id, { scrape_status: 'no_link' });
} }
// Appel Mistral Nemo // Appel LLM (Bifrost)
const enriched = await callMistralWithRetry(row, scrapeContent); const enriched = useNetwork
? await callBifrostWithRetry(row, scrapeData, autresLiens)
: MOCK_BIFROST_RESULT;
if (!enriched) { if (!enriched) {
log(`Échec Mistral sur fiche ${row.Id}, flag ai_error`); log(`Échec Bifrost sur fiche ${row.Id}, flag ai_error`);
await nocodbPatch(NOCODB_TABLE_ORGAS, row.Id, { await patchRow(row.Id, { moderation_status: 'ai_error', ai_processed: true });
moderation_status: 'ai_error',
ai_processed: true
});
continue; continue;
} }
// Normalisation tags // Normalisation tags
const rawTags = enriched.tags_fonction || []; const rawTags = enriched.tags || enriched.tags_fonction || [];
const normalizedTags = [...new Set(rawTags.map(normalizeTag).filter(Boolean))]; const normalizedTags = [...new Set(rawTags.map(normalizeTag).filter(Boolean))];
// Update NocoDB
const updateData = { const updateData = {
description_enrichie: enriched.description_enrichie || null, description_enrichie: enriched.description || null,
points_cles: enriched.points_cles ? JSON.stringify(enriched.points_cles) : null,
tags_fonction: normalizedTags.join(','), tags_fonction: normalizedTags.join(','),
moderation_status: 'ai_processed', moderation_status: 'ai_processed',
ai_processed: true, ai_processed: true,
ai_raw_output: JSON.stringify({ output: enriched, confiance: enriched.confiance }) ai_raw_output: JSON.stringify({ output: enriched, confiance: enriched.confiance }),
}; };
// Conserver echelle/territoire/localisation si l'IA les a enrichis // Nom : ne remplacer que le placeholder posé par le formulaire assoupli (B5-M1).
if (enriched.echelle && !row.echelle) updateData.echelle = enriched.echelle; if (enriched.nom && typeof row.nom === 'string' && row.nom.startsWith('[à qualifier]')) {
if (enriched.territoire && !row.territoire) updateData.territoire = enriched.territoire; updateData.nom = enriched.nom;
if (enriched.localisation_ville && !row.localisation_ville) { }
updateData.localisation_ville = enriched.localisation_ville; if (enriched.ville && !row.localisation_ville) {
updateData.localisation_ville = enriched.ville;
}
// submission_type : ne réajuster que si l'utilisateur n'avait pas choisi de chip
// (marqueur posé par utils/submitLibre.ts::toOrgaPayload — "Type : non précisé").
const typeNonPrecise = typeof row.description_user === 'string'
&& row.description_user.startsWith('Type : non précisé');
if (typeNonPrecise && VALID_SUBMISSION_TYPES.includes(enriched.type_suggere)) {
updateData.submission_type = enriched.type_suggere;
} }
await nocodbPatch(NOCODB_TABLE_ORGAS, row.Id, updateData); await patchRow(row.Id, updateData);
await logUsage(enriched._usage, WORKER_MODEL, 'enrichissement', row.Id);
// Log usage tokens await notifyNtfy(
await logUsage(enriched._usage, 'open-mistral-nemo', 'enrichissement', row.Id); '[AEP] Fiche enrichie',
[
`Id NocoDB : ${row.Id}`,
`Nom suggéré : ${updateData.nom || row.nom}`,
`Type : ${updateData.submission_type || row.submission_type || 'nc'}`,
`Confiance : ${enriched.confiance || 'nc'}`,
].join('\n'),
);
const elapsed = ((Date.now() - rowStart) / 1000).toFixed(1); const elapsed = ((Date.now() - rowStart) / 1000).toFixed(1);
log(`Fiche ${row.Id} traitée en ${elapsed}s — confiance: ${enriched.confiance || 'nc'}`); log(`Fiche ${row.Id} traitée en ${elapsed}s — confiance: ${enriched.confiance || 'nc'}`);
@@ -480,15 +585,12 @@ async function run() {
log(`=== Run terminé: ${processedCount}/${rows.length} fiches traitées en ${((Date.now() - startTime) / 1000).toFixed(1)}s ===`); log(`=== Run terminé: ${processedCount}/${rows.length} fiches traitées en ${((Date.now() - startTime) / 1000).toFixed(1)}s ===`);
// Check seuil modération
await checkModerationQueue();
} catch (e) { } catch (e) {
log('ERREUR WORKER:', e.message); log('ERREUR WORKER:', e.message);
console.error(e.stack); console.error(e.stack);
process.exit(1); process.exitCode = 1;
} finally { } finally {
releaseLock(); if (!DRY_RUN) releaseLock();
} }
} }
+26
View File
@@ -0,0 +1,26 @@
[
{
"Id": 9001,
"nom": "[à qualifier] exemple.org",
"url": "https://exemple.org",
"description_user": "Ça m'a beaucoup aidé à sortir de l'isolement quand j'ai monté mon agence.\n\nLiens :\nhttps://exemple.org\nhttps://exemple.org/guide-installation",
"submitted_by_email": null,
"submission_type": "ecosysteme",
"moderation_status": "pending",
"ai_processed": false,
"scrape_status": "pending",
"localisation_ville": null
},
{
"Id": 9002,
"nom": "[à qualifier] sans titre",
"url": null,
"description_user": "Type : non précisé\n\nUn super groupe d'entraide entre archis installés en rural, très actif sur les questions de reconversion.",
"submitted_by_email": "test@exemple.fr",
"submission_type": "ecosysteme",
"moderation_status": "pending",
"ai_processed": false,
"scrape_status": "no_link",
"localisation_ville": null
}
]