<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>https://infocepo.com/wiki/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Tcepo</id>
	<title>Essential - User contributions [en]</title>
	<link rel="self" type="application/atom+xml" href="https://infocepo.com/wiki/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Tcepo"/>
	<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php/Special:Contributions/Tcepo"/>
	<updated>2026-08-12T07:49:32Z</updated>
	<subtitle>User contributions</subtitle>
	<generator>MediaWiki 1.39.13</generator>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2102</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2102"/>
		<updated>2026-08-12T06:42:46Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Weekly GitHub Trending: ajout de projets pertinents&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Découvrir le cloud et l’IA sur infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''cloud, IA, automatisation et laboratoires''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Vue d’ensemble de l’architecture''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel ; les modèles insuffisants sont remplacés automatiquement et les tâches sans valeur sont supprimées. L’autonomie repose sur des mécanismes techniques qui restent opérationnels sans intervention continue, plutôt que sur de simples règles documentées.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com/docs '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Nouveautés du 02/06 au 31/07/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] : API de traduction multilingue temps réel — PR #3 fix timeout (gestion réponses lentes), PR #2 URLs configurables via .env.&lt;br /&gt;
* [https://github.com/ynotopec/GLM-4.7-Flash-vLLM '''GLM-4.7-Flash-vLLM'''] : Déploiement vLLM GLM-4.7 Flash.&lt;br /&gt;
* [https://github.com/ynotopec/qwen3.5-fp8-vllm '''qwen3.5-fp8-vllm'''] : Déploiement Qwen 3.5 FP8 via vLLM.&lt;br /&gt;
* [https://github.com/ynotopec/api-rag '''api-rag'''] : API RAG (recherche sémantique) — fix support FAISS SVE, suppression logs verbeux, wait-for-APIS avant build index vectoriel.&lt;br /&gt;
* [https://github.com/ynotopec/api-diarization '''api-diarization'''] : API de diarisation audio — installation améliorée, ergonomie de l’authentification et documentation de configuration du projet.&lt;br /&gt;
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] : API FastAPI ASR (Whisper) — refonte complète, installation avec uv et documentation d’exécution.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-openmed '''privacy-proxy-openmed'''] : Privacy proxy OpenMed-based — LLM upstream optionnel (mode privacy-only).&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''privacy-proxy-gliner2'''] : Privacy proxy GLiNER2-based — détection entités nommées, upstream LLM optionnel, PR #1 en cours.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] : Proxy LLM personnalisé — paramètres de modèle configurables, alias &amp;lt;code&amp;gt;ai-default&amp;lt;/code&amp;gt; et gestion des erreurs amont.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Priorités &amp;amp; Veille ==&lt;br /&gt;
&lt;br /&gt;
Les priorités ci-dessous sont les seules entrées de veille conservées. Chaque item doit avoir un '''statut''' : ''en cours'', ''testé'', ''retenu'' ou ''abandonné''. Les projets sans action réelle sont retirés à chaque révision.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] : routage sémantique de requêtes vers les bons modèles.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache'''] : couche KV cache pour réduire la latence d'inférence.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] : orchestration de workflows critiques avec persistence.&lt;br /&gt;
* [https://github.com/lyogavin/airllm '''airLLM'''] : inférence LLM sur GPU de 4 Go — pour réduire les coûts d’inférence locale.&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute'''] : passerelle multi-providers (50+ gratuits), compression 15-95% de tokens — à évaluer comme fallback.&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistant IA auto-hébergé'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Assistants IA &amp;amp; Agents ===&lt;br /&gt;
* [https://github.com/TencentCloud/TencentDB-Agent-Memory '''TencentDB-Agent-Memory'''] : hub de mémoire d’équipe pour agents IA — transforme conversations, documents et code en ressources réutilisables.&lt;br /&gt;
* [https://github.com/unclebob/swarm-forge '''Swarm Forge'''] : Outil simple de coordination de plusieurs agents IA.&lt;br /&gt;
* [https://github.com/embabel/embabel-agent '''Embabel Agent'''] : Framework agent pour JVM (Kotlin/Java).&lt;br /&gt;
* [https://github.com/zhaoxuya520/reverse-skill '''reverse-skill'''] : ensemble de skills de routage pour rétro-ingénierie et pentest — routage IA et initialisation automatique de la chaîne d’outils.&lt;br /&gt;
* [https://github.com/esengine/DeepSeek-Reasonix '''DeepSeek-Reasonix'''] : Agent de codage terminal natif DeepSeek, optimisé pour la stabilité du prefix-cache.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd'''] : skill qui incite l’agent de code à fournir une réponse claire et orientée action.&lt;br /&gt;
* [https://github.com/block/buzz '''Buzz (Block)'''] : Plateforme de communication type hive mind pour collaboration multi-agents.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/livekit/agents '''LiveKit Agents'''] : Framework Python/Node.js pour agents IA vocaux en temps réel — build, deployer et scaler des agents multimodaux en production. (1 100+ ★)&lt;br /&gt;
* [https://github.com/huangruiteng/loopx '''loopx'''] : noyau d’ingénierie de boucle légère pour les équipes d’agents IA à long terme, agnostique aux agents de code, avec objectifs durables et transferts vérifiables.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/cloudflare/computer '''cloudflare/computer'''] : environnement d’exécution pour agents IA — terminal complet et système d’exploitation pour exécuter des agents autonomes.&lt;br /&gt;
* [https://github.com/google/skills '''google/skills'''] : ensemble de skills IA pour les produits Google — intégration d’outils et de services Google pour les agents de code.&lt;br /&gt;
* [https://github.com/uber/adr '''adr'''] - ADR secures enterprise AI agents through observability, security benchmarking, and threat detection. Deployed at Uber..&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] : transforme tout PDF technique en skill Claude Code, prêt à être étudié et référencé.&lt;br /&gt;
* [https://github.com/firecrawl/pdf-inspector '''pdf-inspector'''] : bibliothèque Rust rapide pour l’inspection, la classification et l’extraction de texte PDF — détection intelligente des scans et du texte pour le routage RAG. (8 600+ ★)&lt;br /&gt;
* [https://github.com/vitali87/code-graph-rag '''Code-Graph-RAG'''] : RAG pour monorepo via graphe de connaissances — interroge, comprend et édite une base de code multilangage avec l’IA et Tree-sitter ; compatible MCP.&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/antirez/ds4 '''ds4'''] : moteur d’inférence local DeepSeek 4 Flash/PRO pour Metal, CUDA et ROCm, par Antirez.&lt;br /&gt;
&lt;br /&gt;
=== Outils complémentaires ===&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] : outil open source de revue de code testé à l’échelle d’Alibaba : pipeline déterministe et agent LLM avec commentaires précis par ligne. (16 590 ★)&lt;br /&gt;
* [https://github.com/agavra/tuicr '''tuicr'''] : outil de revue de code en TUI avec raccourcis Vim.&lt;br /&gt;
* [https://github.com/goauthentik/authentik '''authentik'''] : fournisseur d’identité open source auto-hébergé — SSO, MFA, SCIM et fédération d’identité ; alternative à Okta/Auth0. (54 000+ ★)&lt;br /&gt;
* [https://github.com/semantica-agi/semantica '''semantica'''] : infrastructure orientée graphe pour le contexte et des systèmes IA responsables.&lt;br /&gt;
* [https://github.com/drawdb-io/drawdb '''drawdb'''] : éditeur en ligne gratuit et intuitif de diagrammes de bases de données.&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D '''GROQ LLM accelerator''']&lt;br /&gt;
* [https://github.com/Comfy-Org/ComfyUI '''ComfyUI'''] : interface à nœuds pour modèles de diffusion — API, backend modulaire et exécuteur de workflows locaux pour la génération d’images et l’IA visuelle. (125 000+ ★)&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Flux direct : code → CI/CD → déploiement → supervision. Chaque étape se déclenche automatiquement lorsque la précédente réussit. En cas d’échec, une alerte est émise directement ; une intervention humaine n’est requise qu’en cas de décision explicite de repli.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 DataLab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont volontairement limitées en journée.&lt;br /&gt;
* Le coût unitaire est mesuré automatiquement dans des tableaux de bord en temps réel.&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser des fichiers, tampons, files d’attente et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2101</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2101"/>
		<updated>2026-08-12T06:41:39Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Rollback: post-edit verification failed&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Découvrir le cloud et l’IA sur infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''cloud, IA, automatisation et laboratoires''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Vue d’ensemble de l’architecture''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel ; les modèles insuffisants sont remplacés automatiquement et les tâches sans valeur sont supprimées. L’autonomie repose sur des mécanismes techniques qui restent opérationnels sans intervention continue, plutôt que sur de simples règles documentées.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com/docs '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Nouveautés du 02/06 au 31/07/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] : API de traduction multilingue temps réel — PR #3 fix timeout (gestion réponses lentes), PR #2 URLs configurables via .env.&lt;br /&gt;
* [https://github.com/ynotopec/GLM-4.7-Flash-vLLM '''GLM-4.7-Flash-vLLM'''] : Déploiement vLLM GLM-4.7 Flash.&lt;br /&gt;
* [https://github.com/ynotopec/qwen3.5-fp8-vllm '''qwen3.5-fp8-vllm'''] : Déploiement Qwen 3.5 FP8 via vLLM.&lt;br /&gt;
* [https://github.com/ynotopec/api-rag '''api-rag'''] : API RAG (recherche sémantique) — fix support FAISS SVE, suppression logs verbeux, wait-for-APIS avant build index vectoriel.&lt;br /&gt;
* [https://github.com/ynotopec/api-diarization '''api-diarization'''] : API de diarisation audio — installation améliorée, ergonomie de l’authentification et documentation de configuration du projet.&lt;br /&gt;
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] : API FastAPI ASR (Whisper) — refonte complète, installation avec uv et documentation d’exécution.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-openmed '''privacy-proxy-openmed'''] : Privacy proxy OpenMed-based — LLM upstream optionnel (mode privacy-only).&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''privacy-proxy-gliner2'''] : Privacy proxy GLiNER2-based — détection entités nommées, upstream LLM optionnel, PR #1 en cours.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] : Proxy LLM personnalisé — paramètres de modèle configurables, alias &amp;lt;code&amp;gt;ai-default&amp;lt;/code&amp;gt; et gestion des erreurs amont.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Priorités &amp;amp; Veille ==&lt;br /&gt;
&lt;br /&gt;
Les priorités ci-dessous sont les seules entrées de veille conservées. Chaque item doit avoir un '''statut''' : ''en cours'', ''testé'', ''retenu'' ou ''abandonné''. Les projets sans action réelle sont retirés à chaque révision.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] : routage sémantique de requêtes vers les bons modèles.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache'''] : couche KV cache pour réduire la latence d'inférence.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] : orchestration de workflows critiques avec persistence.&lt;br /&gt;
* [https://github.com/lyogavin/airllm '''airLLM'''] : inférence LLM sur GPU de 4 Go — pour réduire les coûts d’inférence locale.&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute'''] : passerelle multi-providers (50+ gratuits), compression 15-95% de tokens — à évaluer comme fallback.&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistant IA auto-hébergé'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Assistants IA &amp;amp; Agents ===&lt;br /&gt;
* [https://github.com/TencentCloud/TencentDB-Agent-Memory '''TencentDB-Agent-Memory'''] : hub de mémoire d’équipe pour agents IA — transforme conversations, documents et code en ressources réutilisables.&lt;br /&gt;
* [https://github.com/unclebob/swarm-forge '''Swarm Forge'''] : Outil simple de coordination de plusieurs agents IA.&lt;br /&gt;
* [https://github.com/embabel/embabel-agent '''Embabel Agent'''] : Framework agent pour JVM (Kotlin/Java).&lt;br /&gt;
* [https://github.com/zhaoxuya520/reverse-skill '''reverse-skill'''] : ensemble de skills de routage pour rétro-ingénierie et pentest — routage IA et initialisation automatique de la chaîne d’outils.&lt;br /&gt;
* [https://github.com/esengine/DeepSeek-Reasonix '''DeepSeek-Reasonix'''] : Agent de codage terminal natif DeepSeek, optimisé pour la stabilité du prefix-cache.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd'''] : skill qui incite l’agent de code à fournir une réponse claire et orientée action.&lt;br /&gt;
* [https://github.com/block/buzz '''Buzz (Block)'''] : Plateforme de communication type hive mind pour collaboration multi-agents.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/livekit/agents '''LiveKit Agents'''] : Framework Python/Node.js pour agents IA vocaux en temps réel — build, deployer et scaler des agents multimodaux en production. (1 100+ ★)&lt;br /&gt;
* [https://github.com/huangruiteng/loopx '''loopx'''] : noyau d’ingénierie de boucle légère pour les équipes d’agents IA à long terme, agnostique aux agents de code, avec objectifs durables et transferts vérifiables.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/cloudflare/computer '''cloudflare/computer'''] : environnement d’exécution pour agents IA — terminal complet et système d’exploitation pour exécuter des agents autonomes.&lt;br /&gt;
* [https://github.com/google/skills '''google/skills'''] : ensemble de skills IA pour les produits Google — intégration d’outils et de services Google pour les agents de code.&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] : transforme tout PDF technique en skill Claude Code, prêt à être étudié et référencé.&lt;br /&gt;
* [https://github.com/firecrawl/pdf-inspector '''pdf-inspector'''] : bibliothèque Rust rapide pour l’inspection, la classification et l’extraction de texte PDF — détection intelligente des scans et du texte pour le routage RAG. (8 600+ ★)&lt;br /&gt;
* [https://github.com/vitali87/code-graph-rag '''Code-Graph-RAG'''] : RAG pour monorepo via graphe de connaissances — interroge, comprend et édite une base de code multilangage avec l’IA et Tree-sitter ; compatible MCP.&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/antirez/ds4 '''ds4'''] : moteur d’inférence local DeepSeek 4 Flash/PRO pour Metal, CUDA et ROCm, par Antirez.&lt;br /&gt;
&lt;br /&gt;
=== Outils complémentaires ===&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] : outil open source de revue de code testé à l’échelle d’Alibaba : pipeline déterministe et agent LLM avec commentaires précis par ligne. (16 590 ★)&lt;br /&gt;
* [https://github.com/agavra/tuicr '''tuicr'''] : outil de revue de code en TUI avec raccourcis Vim.&lt;br /&gt;
* [https://github.com/goauthentik/authentik '''authentik'''] : fournisseur d’identité open source auto-hébergé — SSO, MFA, SCIM et fédération d’identité ; alternative à Okta/Auth0. (54 000+ ★)&lt;br /&gt;
* [https://github.com/semantica-agi/semantica '''semantica'''] : infrastructure orientée graphe pour le contexte et des systèmes IA responsables.&lt;br /&gt;
* [https://github.com/drawdb-io/drawdb '''drawdb'''] : éditeur en ligne gratuit et intuitif de diagrammes de bases de données.&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D '''GROQ LLM accelerator''']&lt;br /&gt;
* [https://github.com/Comfy-Org/ComfyUI '''ComfyUI'''] : interface à nœuds pour modèles de diffusion — API, backend modulaire et exécuteur de workflows locaux pour la génération d’images et l’IA visuelle. (125 000+ ★)&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Flux direct : code → CI/CD → déploiement → supervision. Chaque étape se déclenche automatiquement lorsque la précédente réussit. En cas d’échec, une alerte est émise directement ; une intervention humaine n’est requise qu’en cas de décision explicite de repli.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 DataLab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont volontairement limitées en journée.&lt;br /&gt;
* Le coût unitaire est mesuré automatiquement dans des tableaux de bord en temps réel.&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser des fichiers, tampons, files d’attente et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2100</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2100"/>
		<updated>2026-08-12T06:41:39Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Weekly GitHub Trending: ajout de projets pertinents&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Découvrir le cloud et l’IA sur infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''cloud, IA, automatisation et laboratoires''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Vue d’ensemble de l’architecture''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel ; les modèles insuffisants sont remplacés automatiquement et les tâches sans valeur sont supprimées. L’autonomie repose sur des mécanismes techniques qui restent opérationnels sans intervention continue, plutôt que sur de simples règles documentées.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com/docs '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Nouveautés du 02/06 au 31/07/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] : API de traduction multilingue temps réel — PR #3 fix timeout (gestion réponses lentes), PR #2 URLs configurables via .env.&lt;br /&gt;
* [https://github.com/ynotopec/GLM-4.7-Flash-vLLM '''GLM-4.7-Flash-vLLM'''] : Déploiement vLLM GLM-4.7 Flash.&lt;br /&gt;
* [https://github.com/ynotopec/qwen3.5-fp8-vllm '''qwen3.5-fp8-vllm'''] : Déploiement Qwen 3.5 FP8 via vLLM.&lt;br /&gt;
* [https://github.com/ynotopec/api-rag '''api-rag'''] : API RAG (recherche sémantique) — fix support FAISS SVE, suppression logs verbeux, wait-for-APIS avant build index vectoriel.&lt;br /&gt;
* [https://github.com/ynotopec/api-diarization '''api-diarization'''] : API de diarisation audio — installation améliorée, ergonomie de l’authentification et documentation de configuration du projet.&lt;br /&gt;
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] : API FastAPI ASR (Whisper) — refonte complète, installation avec uv et documentation d’exécution.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-openmed '''privacy-proxy-openmed'''] : Privacy proxy OpenMed-based — LLM upstream optionnel (mode privacy-only).&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''privacy-proxy-gliner2'''] : Privacy proxy GLiNER2-based — détection entités nommées, upstream LLM optionnel, PR #1 en cours.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] : Proxy LLM personnalisé — paramètres de modèle configurables, alias &amp;lt;code&amp;gt;ai-default&amp;lt;/code&amp;gt; et gestion des erreurs amont.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Priorités &amp;amp; Veille ==&lt;br /&gt;
&lt;br /&gt;
Les priorités ci-dessous sont les seules entrées de veille conservées. Chaque item doit avoir un '''statut''' : ''en cours'', ''testé'', ''retenu'' ou ''abandonné''. Les projets sans action réelle sont retirés à chaque révision.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] : routage sémantique de requêtes vers les bons modèles.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache'''] : couche KV cache pour réduire la latence d'inférence.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] : orchestration de workflows critiques avec persistence.&lt;br /&gt;
* [https://github.com/lyogavin/airllm '''airLLM'''] : inférence LLM sur GPU de 4 Go — pour réduire les coûts d’inférence locale.&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute'''] : passerelle multi-providers (50+ gratuits), compression 15-95% de tokens — à évaluer comme fallback.&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistant IA auto-hébergé'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Assistants IA &amp;amp; Agents ===&lt;br /&gt;
* [https://github.com/TencentCloud/TencentDB-Agent-Memory '''TencentDB-Agent-Memory'''] : hub de mémoire d’équipe pour agents IA — transforme conversations, documents et code en ressources réutilisables.&lt;br /&gt;
* [https://github.com/unclebob/swarm-forge '''Swarm Forge'''] : Outil simple de coordination de plusieurs agents IA.&lt;br /&gt;
* [https://github.com/embabel/embabel-agent '''Embabel Agent'''] : Framework agent pour JVM (Kotlin/Java).&lt;br /&gt;
* [https://github.com/zhaoxuya520/reverse-skill '''reverse-skill'''] : ensemble de skills de routage pour rétro-ingénierie et pentest — routage IA et initialisation automatique de la chaîne d’outils.&lt;br /&gt;
* [https://github.com/esengine/DeepSeek-Reasonix '''DeepSeek-Reasonix'''] : Agent de codage terminal natif DeepSeek, optimisé pour la stabilité du prefix-cache.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd'''] : skill qui incite l’agent de code à fournir une réponse claire et orientée action.&lt;br /&gt;
* [https://github.com/block/buzz '''Buzz (Block)'''] : Plateforme de communication type hive mind pour collaboration multi-agents.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/livekit/agents '''LiveKit Agents'''] : Framework Python/Node.js pour agents IA vocaux en temps réel — build, deployer et scaler des agents multimodaux en production. (1 100+ ★)&lt;br /&gt;
* [https://github.com/huangruiteng/loopx '''loopx'''] : noyau d’ingénierie de boucle légère pour les équipes d’agents IA à long terme, agnostique aux agents de code, avec objectifs durables et transferts vérifiables.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/cloudflare/computer '''cloudflare/computer'''] : environnement d’exécution pour agents IA — terminal complet et système d’exploitation pour exécuter des agents autonomes.&lt;br /&gt;
* [https://github.com/google/skills '''google/skills'''] : ensemble de skills IA pour les produits Google — intégration d’outils et de services Google pour les agents de code.&lt;br /&gt;
* [https://github.com/uber/adr '''adr'''] - ADR secures enterprise AI agents through observability, security benchmarking, and threat detection. Deployed at Uber..&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] : transforme tout PDF technique en skill Claude Code, prêt à être étudié et référencé.&lt;br /&gt;
* [https://github.com/firecrawl/pdf-inspector '''pdf-inspector'''] : bibliothèque Rust rapide pour l’inspection, la classification et l’extraction de texte PDF — détection intelligente des scans et du texte pour le routage RAG. (8 600+ ★)&lt;br /&gt;
* [https://github.com/vitali87/code-graph-rag '''Code-Graph-RAG'''] : RAG pour monorepo via graphe de connaissances — interroge, comprend et édite une base de code multilangage avec l’IA et Tree-sitter ; compatible MCP.&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/antirez/ds4 '''ds4'''] : moteur d’inférence local DeepSeek 4 Flash/PRO pour Metal, CUDA et ROCm, par Antirez.&lt;br /&gt;
&lt;br /&gt;
=== Outils complémentaires ===&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] : outil open source de revue de code testé à l’échelle d’Alibaba : pipeline déterministe et agent LLM avec commentaires précis par ligne. (16 590 ★)&lt;br /&gt;
* [https://github.com/agavra/tuicr '''tuicr'''] : outil de revue de code en TUI avec raccourcis Vim.&lt;br /&gt;
* [https://github.com/goauthentik/authentik '''authentik'''] : fournisseur d’identité open source auto-hébergé — SSO, MFA, SCIM et fédération d’identité ; alternative à Okta/Auth0. (54 000+ ★)&lt;br /&gt;
* [https://github.com/semantica-agi/semantica '''semantica'''] : infrastructure orientée graphe pour le contexte et des systèmes IA responsables.&lt;br /&gt;
* [https://github.com/drawdb-io/drawdb '''drawdb'''] : éditeur en ligne gratuit et intuitif de diagrammes de bases de données.&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D '''GROQ LLM accelerator''']&lt;br /&gt;
* [https://github.com/Comfy-Org/ComfyUI '''ComfyUI'''] : interface à nœuds pour modèles de diffusion — API, backend modulaire et exécuteur de workflows locaux pour la génération d’images et l’IA visuelle. (125 000+ ★)&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Flux direct : code → CI/CD → déploiement → supervision. Chaque étape se déclenche automatiquement lorsque la précédente réussit. En cas d’échec, une alerte est émise directement ; une intervention humaine n’est requise qu’en cas de décision explicite de repli.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 DataLab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont volontairement limitées en journée.&lt;br /&gt;
* Le coût unitaire est mesuré automatiquement dans des tableaux de bord en temps réel.&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser des fichiers, tampons, files d’attente et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2099</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2099"/>
		<updated>2026-08-11T21:55:46Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Découvrir le cloud et l’IA sur infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''cloud, IA, automatisation et laboratoires''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Vue d’ensemble de l’architecture''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel ; les modèles insuffisants sont remplacés automatiquement et les tâches sans valeur sont supprimées. L’autonomie repose sur des mécanismes techniques qui restent opérationnels sans intervention continue, plutôt que sur de simples règles documentées.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com/docs '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Nouveautés du 02/06 au 31/07/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] : API de traduction multilingue temps réel — PR #3 fix timeout (gestion réponses lentes), PR #2 URLs configurables via .env.&lt;br /&gt;
* [https://github.com/ynotopec/GLM-4.7-Flash-vLLM '''GLM-4.7-Flash-vLLM'''] : Déploiement vLLM GLM-4.7 Flash.&lt;br /&gt;
* [https://github.com/ynotopec/qwen3.5-fp8-vllm '''qwen3.5-fp8-vllm'''] : Déploiement Qwen 3.5 FP8 via vLLM.&lt;br /&gt;
* [https://github.com/ynotopec/api-rag '''api-rag'''] : API RAG (recherche sémantique) — fix support FAISS SVE, suppression logs verbeux, wait-for-APIS avant build index vectoriel.&lt;br /&gt;
* [https://github.com/ynotopec/api-diarization '''api-diarization'''] : API de diarisation audio — installation améliorée, ergonomie de l’authentification et documentation de configuration du projet.&lt;br /&gt;
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] : API FastAPI ASR (Whisper) — refonte complète, installation avec uv et documentation d’exécution.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-openmed '''privacy-proxy-openmed'''] : Privacy proxy OpenMed-based — LLM upstream optionnel (mode privacy-only).&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''privacy-proxy-gliner2'''] : Privacy proxy GLiNER2-based — détection entités nommées, upstream LLM optionnel, PR #1 en cours.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] : Proxy LLM personnalisé — paramètres de modèle configurables, alias &amp;lt;code&amp;gt;ai-default&amp;lt;/code&amp;gt; et gestion des erreurs amont.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Priorités &amp;amp; Veille ==&lt;br /&gt;
&lt;br /&gt;
Les priorités ci-dessous sont les seules entrées de veille conservées. Chaque item doit avoir un '''statut''' : ''en cours'', ''testé'', ''retenu'' ou ''abandonné''. Les projets sans action réelle sont retirés à chaque révision.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] : routage sémantique de requêtes vers les bons modèles.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache'''] : couche KV cache pour réduire la latence d'inférence.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] : orchestration de workflows critiques avec persistence.&lt;br /&gt;
* [https://github.com/lyogavin/airllm '''airLLM'''] : inférence LLM sur GPU de 4 Go — pour réduire les coûts d’inférence locale.&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute'''] : passerelle multi-providers (50+ gratuits), compression 15-95% de tokens — à évaluer comme fallback.&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistant IA auto-hébergé'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Assistants IA &amp;amp; Agents ===&lt;br /&gt;
* [https://github.com/TencentCloud/TencentDB-Agent-Memory '''TencentDB-Agent-Memory'''] : hub de mémoire d’équipe pour agents IA — transforme conversations, documents et code en ressources réutilisables.&lt;br /&gt;
* [https://github.com/unclebob/swarm-forge '''Swarm Forge'''] : Outil simple de coordination de plusieurs agents IA.&lt;br /&gt;
* [https://github.com/embabel/embabel-agent '''Embabel Agent'''] : Framework agent pour JVM (Kotlin/Java).&lt;br /&gt;
* [https://github.com/zhaoxuya520/reverse-skill '''reverse-skill'''] : ensemble de skills de routage pour rétro-ingénierie et pentest — routage IA et initialisation automatique de la chaîne d’outils.&lt;br /&gt;
* [https://github.com/esengine/DeepSeek-Reasonix '''DeepSeek-Reasonix'''] : Agent de codage terminal natif DeepSeek, optimisé pour la stabilité du prefix-cache.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd'''] : skill qui incite l’agent de code à fournir une réponse claire et orientée action.&lt;br /&gt;
* [https://github.com/block/buzz '''Buzz (Block)'''] : Plateforme de communication type hive mind pour collaboration multi-agents.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/livekit/agents '''LiveKit Agents'''] : Framework Python/Node.js pour agents IA vocaux en temps réel — build, deployer et scaler des agents multimodaux en production. (1 100+ ★)&lt;br /&gt;
* [https://github.com/huangruiteng/loopx '''loopx'''] : noyau d’ingénierie de boucle légère pour les équipes d’agents IA à long terme, agnostique aux agents de code, avec objectifs durables et transferts vérifiables.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/cloudflare/computer '''cloudflare/computer'''] : environnement d’exécution pour agents IA — terminal complet et système d’exploitation pour exécuter des agents autonomes.&lt;br /&gt;
* [https://github.com/google/skills '''google/skills'''] : ensemble de skills IA pour les produits Google — intégration d’outils et de services Google pour les agents de code.&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] : transforme tout PDF technique en skill Claude Code, prêt à être étudié et référencé.&lt;br /&gt;
* [https://github.com/firecrawl/pdf-inspector '''pdf-inspector'''] : bibliothèque Rust rapide pour l’inspection, la classification et l’extraction de texte PDF — détection intelligente des scans et du texte pour le routage RAG. (8 600+ ★)&lt;br /&gt;
* [https://github.com/vitali87/code-graph-rag '''Code-Graph-RAG'''] : RAG pour monorepo via graphe de connaissances — interroge, comprend et édite une base de code multilangage avec l’IA et Tree-sitter ; compatible MCP.&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/antirez/ds4 '''ds4'''] : moteur d’inférence local DeepSeek 4 Flash/PRO pour Metal, CUDA et ROCm, par Antirez.&lt;br /&gt;
&lt;br /&gt;
=== Outils complémentaires ===&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] : outil open source de revue de code testé à l’échelle d’Alibaba : pipeline déterministe et agent LLM avec commentaires précis par ligne. (16 590 ★)&lt;br /&gt;
* [https://github.com/agavra/tuicr '''tuicr'''] : outil de revue de code en TUI avec raccourcis Vim.&lt;br /&gt;
* [https://github.com/goauthentik/authentik '''authentik'''] : fournisseur d’identité open source auto-hébergé — SSO, MFA, SCIM et fédération d’identité ; alternative à Okta/Auth0. (54 000+ ★)&lt;br /&gt;
* [https://github.com/semantica-agi/semantica '''semantica'''] : infrastructure orientée graphe pour le contexte et des systèmes IA responsables.&lt;br /&gt;
* [https://github.com/drawdb-io/drawdb '''drawdb'''] : éditeur en ligne gratuit et intuitif de diagrammes de bases de données.&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D '''GROQ LLM accelerator''']&lt;br /&gt;
* [https://github.com/Comfy-Org/ComfyUI '''ComfyUI'''] : interface à nœuds pour modèles de diffusion — API, backend modulaire et exécuteur de workflows locaux pour la génération d’images et l’IA visuelle. (125 000+ ★)&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Flux direct : code → CI/CD → déploiement → supervision. Chaque étape se déclenche automatiquement lorsque la précédente réussit. En cas d’échec, une alerte est émise directement ; une intervention humaine n’est requise qu’en cas de décision explicite de repli.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 DataLab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont volontairement limitées en journée.&lt;br /&gt;
* Le coût unitaire est mesuré automatiquement dans des tableaux de bord en temps réel.&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser des fichiers, tampons, files d’attente et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2098</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2098"/>
		<updated>2026-08-11T21:13:46Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: /* Priorités &amp;amp; Veille */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com/docs '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 02/06 - 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] : API de traduction multilingue temps réel — PR #3 fix timeout (gestion réponses lentes), PR #2 URLs configurables via .env.&lt;br /&gt;
* [https://github.com/ynotopec/GLM-4.7-Flash-vLLM '''GLM-4.7-Flash-vLLM'''] : Déploiement vLLM GLM-4.7 Flash.&lt;br /&gt;
* [https://github.com/ynotopec/qwen3.5-fp8-vllm '''qwen3.5-fp8-vllm'''] : Déploiement Qwen 3.5 FP8 via vLLM.&lt;br /&gt;
* [https://github.com/ynotopec/api-rag '''api-rag'''] : API RAG (recherche sémantique) — fix support FAISS SVE, suppression logs verbeux, wait-for-APIS avant build index vectoriel.&lt;br /&gt;
* [https://github.com/ynotopec/api-diarization '''api-diarization'''] : API diarisation audio — install améliorée, ergonomics auth, documentation setup projet.&lt;br /&gt;
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] : API FastAPI ASR ( Whisper ) — refonte complète, installer uv, docs runtime.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-openmed '''privacy-proxy-openmed'''] : Privacy proxy OpenMed-based — LLM upstream optionnel (mode privacy-only).&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''privacy-proxy-gliner2'''] : Privacy proxy GLiNER2-based — détection entités nommées, upstream LLM optionnel, PR #1 en cours.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] : Proxy LLM custom — paramètres modèle configurables, alias ai-default, gestion erreurs upstream.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Priorités &amp;amp; Veille ==&lt;br /&gt;
&lt;br /&gt;
Les priorités ci-dessous sont les seules entrées de veille conservées. Chaque item doit avoir un **statu** : *en cours*, *testé*, *retenu*, *abandonné*. Les projets listés sans action réelle sont retirés à chaque révision.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] : routage sémantique de requêtes vers les bons modèles.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache'''] : couche KV cache pour réduire la latence d'inférence.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] : orchestration de workflows critiques avec persistence.&lt;br /&gt;
* [https://github.com/lyogavin/airllm '''airLLM'''] : inférence LLM sur GPU 4 Go — pour réduire les coûts de inférence locale.&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute'''] : passerelle multi-providers (50+ gratuits), compression 15-95% de tokens — à évaluer comme fallback.&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Assistants IA &amp;amp; Agents ===&lt;br /&gt;
* [https://github.com/TencentCloud/TencentDB-Agent-Memory '''TencentDB-Agent-Memory'''] : Hub de mémoire d équipe pour agents IA - transforme conversations, docs et code en assets réutilisables.&lt;br /&gt;
* [https://github.com/unclebob/swarm-forge '''Swarm Forge'''] : Outil simple de coordination de plusieurs agents IA.&lt;br /&gt;
* [https://github.com/embabel/embabel-agent '''Embabel Agent'''] : Framework agent pour JVM (Kotlin/Java).&lt;br /&gt;
* [https://github.com/zhaoxuya520/reverse-skill '''reverse-skill'''] : Skill Router Pack pour reverse engineering et pentesting - routage IA, bootstrapping toolchain auto.&lt;br /&gt;
* [https://github.com/esengine/DeepSeek-Reasonix '''DeepSeek-Reasonix'''] : Agent de codage terminal natif DeepSeek, optimisé pour la stabilité du prefix-cache.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd'''] : Skill qui empêche l agent de coder de cacher la réponse - output clair et orienté action.&lt;br /&gt;
* [https://github.com/block/buzz '''Buzz (Block)'''] : Plateforme de communication type hive mind pour collaboration multi-agents.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/livekit/agents '''LiveKit Agents'''] : Framework Python/Node.js pour agents IA vocaux en temps réel — build, deployer et scaler des agents multimodaux en production. (1 100+ ★)&lt;br /&gt;
* [https://github.com/huangruiteng/loopx '''loopx''' ] - Noyau d'ingénierie de boucle légère pour équipes d'agents IA à long terme. Agnostique aux agents de codage, avec objectifs durables et transferts vérifiables.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/cloudflare/computer '''cloudflare/computer'''] - Environnement d'exécution pour agents IA — un terminal complet et un système d'exploitation pour exécuter des agents autonomes.&lt;br /&gt;
* [https://github.com/google/skills '''google/skills'''] - Ensemble de skills IA pour les produits Google — intégration d'outils et services Google pour les agents de codage.&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] : Transforme n importe quel PDF technique en skill Claude Code - prêt à étudier et référencer.&lt;br /&gt;
* [https://github.com/firecrawl/pdf-inspector '''pdf-inspector'''] : Bibliothèque Rust ultra-rapide pour l'inspection, la classification et l'extraction de texte PDF — detection intelligente scans vs texte pour le routing RAG. (8 600+ ★)&lt;br /&gt;
* [https://github.com/vitali87/code-graph-rag '''Code-Graph-RAG'''] : RAG pour monorepo via graphe de connaissances — interrogez, comprenez et editez votre codebase multi-langage avec l'IA et Tree-sitter. MCP server compatible.&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/antirez/ds4 '''ds4'''] : Moteur inference local DeepSeek 4 Flash/PRO pour Metal, CUDA et ROCm - par Antirez.&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D '''GROQ LLM accelerator''']&lt;br /&gt;
* [https://github.com/Comfy-Org/ComfyUI '''ComfyUI''' ] : Interface nodes pour modèles de diffusion — API, backend modular et exécuteur de workflows locaux pour génération d'images et IA visuelle. (125 000+ ★)&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
* [https://github.com/agavra/tuicr '''tuicr''' ] - Revue de code TUI avec raccourcis vim.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/goauthentik/authentik '''authentik'''] : Fournisseur d'identite open-source auto-herberge — SSO, MFA, SCIM, federation d'identite. L'alternative open-source a Okta/Auth0. (54k+ ★)&lt;br /&gt;
* [https://github.com/semantica-agi/semantica '''semantica''' ] - Infrastructure graph-native pour contexte et systèmes IA responsables.&lt;br /&gt;
* [https://github.com/drawdb-io/drawdb '''drawdb''' ] - Éditeur de diagrammes de base de données en ligne gratuit et intuitif.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2097</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2097"/>
		<updated>2026-08-11T21:13:05Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: /* Top tasks */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com/docs '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 02/06 - 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] : API de traduction multilingue temps réel — PR #3 fix timeout (gestion réponses lentes), PR #2 URLs configurables via .env.&lt;br /&gt;
* [https://github.com/ynotopec/GLM-4.7-Flash-vLLM '''GLM-4.7-Flash-vLLM'''] : Déploiement vLLM GLM-4.7 Flash.&lt;br /&gt;
* [https://github.com/ynotopec/qwen3.5-fp8-vllm '''qwen3.5-fp8-vllm'''] : Déploiement Qwen 3.5 FP8 via vLLM.&lt;br /&gt;
* [https://github.com/ynotopec/api-rag '''api-rag'''] : API RAG (recherche sémantique) — fix support FAISS SVE, suppression logs verbeux, wait-for-APIS avant build index vectoriel.&lt;br /&gt;
* [https://github.com/ynotopec/api-diarization '''api-diarization'''] : API diarisation audio — install améliorée, ergonomics auth, documentation setup projet.&lt;br /&gt;
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] : API FastAPI ASR ( Whisper ) — refonte complète, installer uv, docs runtime.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-openmed '''privacy-proxy-openmed'''] : Privacy proxy OpenMed-based — LLM upstream optionnel (mode privacy-only).&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''privacy-proxy-gliner2'''] : Privacy proxy GLiNER2-based — détection entités nommées, upstream LLM optionnel, PR #1 en cours.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] : Proxy LLM custom — paramètres modèle configurables, alias ai-default, gestion erreurs upstream.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Priorités &amp;amp; Veille ==&lt;br /&gt;
&lt;br /&gt;
Les priorités ci-dessous sont les seules entrées de veille conservées. Chaque item doit avoir un **statu** : *en cours*, *testé*, *retenu*, *abandonné*. Les projets listés sans action réelle sont retirés à chaque révision.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration d'inférence multi-nœuds.&lt;br /&gt;
* [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning et benchmark réaliste.&lt;br /&gt;
* [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques pour les LLM.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix / token en temps réel.&lt;br /&gt;
* '''Qualité auto''' : précision / taux d'hallucination sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] : routage sémantique de requêtes vers les bons modèles.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache'''] : couche KV cache pour réduire la latence d'inférence.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] : orchestration de workflows critiques avec persistence.&lt;br /&gt;
* [https://github.com/lyogavin/airllm '''airLLM'''] : inférence LLM sur GPU 4 Go — pour réduire les coûts de inférence locale.&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute'''] : passerelle multi-providers (50+ gratuits), compression 15-95% de tokens — à évaluer comme fallback.&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Assistants IA &amp;amp; Agents ===&lt;br /&gt;
* [https://github.com/TencentCloud/TencentDB-Agent-Memory '''TencentDB-Agent-Memory'''] : Hub de mémoire d équipe pour agents IA - transforme conversations, docs et code en assets réutilisables.&lt;br /&gt;
* [https://github.com/unclebob/swarm-forge '''Swarm Forge'''] : Outil simple de coordination de plusieurs agents IA.&lt;br /&gt;
* [https://github.com/embabel/embabel-agent '''Embabel Agent'''] : Framework agent pour JVM (Kotlin/Java).&lt;br /&gt;
* [https://github.com/zhaoxuya520/reverse-skill '''reverse-skill'''] : Skill Router Pack pour reverse engineering et pentesting - routage IA, bootstrapping toolchain auto.&lt;br /&gt;
* [https://github.com/esengine/DeepSeek-Reasonix '''DeepSeek-Reasonix'''] : Agent de codage terminal natif DeepSeek, optimisé pour la stabilité du prefix-cache.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd'''] : Skill qui empêche l agent de coder de cacher la réponse - output clair et orienté action.&lt;br /&gt;
* [https://github.com/block/buzz '''Buzz (Block)'''] : Plateforme de communication type hive mind pour collaboration multi-agents.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/livekit/agents '''LiveKit Agents'''] : Framework Python/Node.js pour agents IA vocaux en temps réel — build, deployer et scaler des agents multimodaux en production. (1 100+ ★)&lt;br /&gt;
* [https://github.com/huangruiteng/loopx '''loopx''' ] - Noyau d'ingénierie de boucle légère pour équipes d'agents IA à long terme. Agnostique aux agents de codage, avec objectifs durables et transferts vérifiables.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/cloudflare/computer '''cloudflare/computer'''] - Environnement d'exécution pour agents IA — un terminal complet et un système d'exploitation pour exécuter des agents autonomes.&lt;br /&gt;
* [https://github.com/google/skills '''google/skills'''] - Ensemble de skills IA pour les produits Google — intégration d'outils et services Google pour les agents de codage.&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] : Transforme n importe quel PDF technique en skill Claude Code - prêt à étudier et référencer.&lt;br /&gt;
* [https://github.com/firecrawl/pdf-inspector '''pdf-inspector'''] : Bibliothèque Rust ultra-rapide pour l'inspection, la classification et l'extraction de texte PDF — detection intelligente scans vs texte pour le routing RAG. (8 600+ ★)&lt;br /&gt;
* [https://github.com/vitali87/code-graph-rag '''Code-Graph-RAG'''] : RAG pour monorepo via graphe de connaissances — interrogez, comprenez et editez votre codebase multi-langage avec l'IA et Tree-sitter. MCP server compatible.&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/antirez/ds4 '''ds4'''] : Moteur inference local DeepSeek 4 Flash/PRO pour Metal, CUDA et ROCm - par Antirez.&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D '''GROQ LLM accelerator''']&lt;br /&gt;
* [https://github.com/Comfy-Org/ComfyUI '''ComfyUI''' ] : Interface nodes pour modèles de diffusion — API, backend modular et exécuteur de workflows locaux pour génération d'images et IA visuelle. (125 000+ ★)&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
* [https://github.com/agavra/tuicr '''tuicr''' ] - Revue de code TUI avec raccourcis vim.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/goauthentik/authentik '''authentik'''] : Fournisseur d'identite open-source auto-herberge — SSO, MFA, SCIM, federation d'identite. L'alternative open-source a Okta/Auth0. (54k+ ★)&lt;br /&gt;
* [https://github.com/semantica-agi/semantica '''semantica''' ] - Infrastructure graph-native pour contexte et systèmes IA responsables.&lt;br /&gt;
* [https://github.com/drawdb-io/drawdb '''drawdb''' ] - Éditeur de diagrammes de base de données en ligne gratuit et intuitif.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2096</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2096"/>
		<updated>2026-08-11T21:10:51Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Weekly GitHub Trending: ajout de projets pertinents&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com/docs '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 02/06 - 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] : API de traduction multilingue temps réel — PR #3 fix timeout (gestion réponses lentes), PR #2 URLs configurables via .env.&lt;br /&gt;
* [https://github.com/ynotopec/GLM-4.7-Flash-vLLM '''GLM-4.7-Flash-vLLM'''] : Déploiement vLLM GLM-4.7 Flash.&lt;br /&gt;
* [https://github.com/ynotopec/qwen3.5-fp8-vllm '''qwen3.5-fp8-vllm'''] : Déploiement Qwen 3.5 FP8 via vLLM.&lt;br /&gt;
* [https://github.com/ynotopec/api-rag '''api-rag'''] : API RAG (recherche sémantique) — fix support FAISS SVE, suppression logs verbeux, wait-for-APIS avant build index vectoriel.&lt;br /&gt;
* [https://github.com/ynotopec/api-diarization '''api-diarization'''] : API diarisation audio — install améliorée, ergonomics auth, documentation setup projet.&lt;br /&gt;
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] : API FastAPI ASR ( Whisper ) — refonte complète, installer uv, docs runtime.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-openmed '''privacy-proxy-openmed'''] : Privacy proxy OpenMed-based — LLM upstream optionnel (mode privacy-only).&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''privacy-proxy-gliner2'''] : Privacy proxy GLiNER2-based — détection entités nommées, upstream LLM optionnel, PR #1 en cours.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] : Proxy LLM custom — paramètres modèle configurables, alias ai-default, gestion erreurs upstream.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Priorités &amp;amp; Veille ==&lt;br /&gt;
&lt;br /&gt;
Les priorités ci-dessous sont les seules entrées de veille conservées. Chaque item doit avoir un **statu** : *en cours*, *testé*, *retenu*, *abandonné*. Les projets listés sans action réelle sont retirés à chaque révision.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration d'inférence multi-nœuds.&lt;br /&gt;
* [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning et benchmark réaliste.&lt;br /&gt;
* [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques pour les LLM.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix / token en temps réel.&lt;br /&gt;
* '''Qualité auto''' : précision / taux d'hallucination sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] : routage sémantique de requêtes vers les bons modèles.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache'''] : couche KV cache pour réduire la latence d'inférence.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] : orchestration de workflows critiques avec persistence.&lt;br /&gt;
* [https://github.com/lyogavin/airllm '''airLLM'''] : inférence LLM sur GPU 4 Go — pour réduire les coûts de inférence locale.&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute'''] : passerelle multi-providers (50+ gratuits), compression 15-95% de tokens — à évaluer comme fallback.&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Assistants IA &amp;amp; Agents ===&lt;br /&gt;
* [https://github.com/TencentCloud/TencentDB-Agent-Memory '''TencentDB-Agent-Memory'''] : Hub de mémoire d équipe pour agents IA - transforme conversations, docs et code en assets réutilisables.&lt;br /&gt;
* [https://github.com/unclebob/swarm-forge '''Swarm Forge'''] : Outil simple de coordination de plusieurs agents IA.&lt;br /&gt;
* [https://github.com/embabel/embabel-agent '''Embabel Agent'''] : Framework agent pour JVM (Kotlin/Java).&lt;br /&gt;
* [https://github.com/zhaoxuya520/reverse-skill '''reverse-skill'''] : Skill Router Pack pour reverse engineering et pentesting - routage IA, bootstrapping toolchain auto.&lt;br /&gt;
* [https://github.com/esengine/DeepSeek-Reasonix '''DeepSeek-Reasonix'''] : Agent de codage terminal natif DeepSeek, optimisé pour la stabilité du prefix-cache.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd'''] : Skill qui empêche l agent de coder de cacher la réponse - output clair et orienté action.&lt;br /&gt;
* [https://github.com/block/buzz '''Buzz (Block)'''] : Plateforme de communication type hive mind pour collaboration multi-agents.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/livekit/agents '''LiveKit Agents'''] : Framework Python/Node.js pour agents IA vocaux en temps réel — build, deployer et scaler des agents multimodaux en production. (1 100+ ★)&lt;br /&gt;
* [https://github.com/huangruiteng/loopx '''loopx''' ] - Noyau d'ingénierie de boucle légère pour équipes d'agents IA à long terme. Agnostique aux agents de codage, avec objectifs durables et transferts vérifiables.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/cloudflare/computer '''cloudflare/computer'''] - Environnement d'exécution pour agents IA — un terminal complet et un système d'exploitation pour exécuter des agents autonomes.&lt;br /&gt;
* [https://github.com/google/skills '''google/skills'''] - Ensemble de skills IA pour les produits Google — intégration d'outils et services Google pour les agents de codage.&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] : Transforme n importe quel PDF technique en skill Claude Code - prêt à étudier et référencer.&lt;br /&gt;
* [https://github.com/firecrawl/pdf-inspector '''pdf-inspector'''] : Bibliothèque Rust ultra-rapide pour l'inspection, la classification et l'extraction de texte PDF — detection intelligente scans vs texte pour le routing RAG. (8 600+ ★)&lt;br /&gt;
* [https://github.com/vitali87/code-graph-rag '''Code-Graph-RAG'''] : RAG pour monorepo via graphe de connaissances — interrogez, comprenez et editez votre codebase multi-langage avec l'IA et Tree-sitter. MCP server compatible.&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/antirez/ds4 '''ds4'''] : Moteur inference local DeepSeek 4 Flash/PRO pour Metal, CUDA et ROCm - par Antirez.&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D '''GROQ LLM accelerator''']&lt;br /&gt;
* [https://github.com/Comfy-Org/ComfyUI '''ComfyUI''' ] : Interface nodes pour modèles de diffusion — API, backend modular et exécuteur de workflows locaux pour génération d'images et IA visuelle. (125 000+ ★)&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
* [https://github.com/agavra/tuicr '''tuicr''' ] - Revue de code TUI avec raccourcis vim.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/goauthentik/authentik '''authentik'''] : Fournisseur d'identite open-source auto-herberge — SSO, MFA, SCIM, federation d'identite. L'alternative open-source a Okta/Auth0. (54k+ ★)&lt;br /&gt;
* [https://github.com/semantica-agi/semantica '''semantica''' ] - Infrastructure graph-native pour contexte et systèmes IA responsables.&lt;br /&gt;
* [https://github.com/drawdb-io/drawdb '''drawdb''' ] - Éditeur de diagrammes de base de données en ligne gratuit et intuitif.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2095</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2095"/>
		<updated>2026-08-11T20:43:29Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: /* Parcours recommandés */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com/docs '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 02/06 - 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] : API de traduction multilingue temps réel — PR #3 fix timeout (gestion réponses lentes), PR #2 URLs configurables via .env.&lt;br /&gt;
* [https://github.com/ynotopec/GLM-4.7-Flash-vLLM '''GLM-4.7-Flash-vLLM'''] : Déploiement vLLM GLM-4.7 Flash.&lt;br /&gt;
* [https://github.com/ynotopec/qwen3.5-fp8-vllm '''qwen3.5-fp8-vllm'''] : Déploiement Qwen 3.5 FP8 via vLLM.&lt;br /&gt;
* [https://github.com/ynotopec/api-rag '''api-rag'''] : API RAG (recherche sémantique) — fix support FAISS SVE, suppression logs verbeux, wait-for-APIS avant build index vectoriel.&lt;br /&gt;
* [https://github.com/ynotopec/api-diarization '''api-diarization'''] : API diarisation audio — install améliorée, ergonomics auth, documentation setup projet.&lt;br /&gt;
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] : API FastAPI ASR ( Whisper ) — refonte complète, installer uv, docs runtime.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-openmed '''privacy-proxy-openmed'''] : Privacy proxy OpenMed-based — LLM upstream optionnel (mode privacy-only).&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''privacy-proxy-gliner2'''] : Privacy proxy GLiNER2-based — détection entités nommées, upstream LLM optionnel, PR #1 en cours.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] : Proxy LLM custom — paramètres modèle configurables, alias ai-default, gestion erreurs upstream.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Priorités &amp;amp; Veille ==&lt;br /&gt;
&lt;br /&gt;
Les priorités ci-dessous sont les seules entrées de veille conservées. Chaque item doit avoir un **statu** : *en cours*, *testé*, *retenu*, *abandonné*. Les projets listés sans action réelle sont retirés à chaque révision.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration d'inférence multi-nœuds.&lt;br /&gt;
* [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning et benchmark réaliste.&lt;br /&gt;
* [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques pour les LLM.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix / token en temps réel.&lt;br /&gt;
* '''Qualité auto''' : précision / taux d'hallucination sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] : routage sémantique de requêtes vers les bons modèles.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache'''] : couche KV cache pour réduire la latence d'inférence.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] : orchestration de workflows critiques avec persistence.&lt;br /&gt;
* [https://github.com/lyogavin/airllm '''airLLM'''] : inférence LLM sur GPU 4 Go — pour réduire les coûts de inférence locale.&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute'''] : passerelle multi-providers (50+ gratuits), compression 15-95% de tokens — à évaluer comme fallback.&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Assistants IA &amp;amp; Agents ===&lt;br /&gt;
* [https://github.com/TencentCloud/TencentDB-Agent-Memory '''TencentDB-Agent-Memory'''] : Hub de mémoire d équipe pour agents IA - transforme conversations, docs et code en assets réutilisables.&lt;br /&gt;
* [https://github.com/unclebob/swarm-forge '''Swarm Forge'''] : Outil simple de coordination de plusieurs agents IA.&lt;br /&gt;
* [https://github.com/embabel/embabel-agent '''Embabel Agent'''] : Framework agent pour JVM (Kotlin/Java).&lt;br /&gt;
* [https://github.com/zhaoxuya520/reverse-skill '''reverse-skill'''] : Skill Router Pack pour reverse engineering et pentesting - routage IA, bootstrapping toolchain auto.&lt;br /&gt;
* [https://github.com/esengine/DeepSeek-Reasonix '''DeepSeek-Reasonix'''] : Agent de codage terminal natif DeepSeek, optimisé pour la stabilité du prefix-cache.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd'''] : Skill qui empêche l agent de coder de cacher la réponse - output clair et orienté action.&lt;br /&gt;
* [https://github.com/block/buzz '''Buzz (Block)'''] : Plateforme de communication type hive mind pour collaboration multi-agents.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/livekit/agents '''LiveKit Agents'''] : Framework Python/Node.js pour agents IA vocaux en temps réel — build, deployer et scaler des agents multimodaux en production. (1 100+ ★)&lt;br /&gt;
* [https://github.com/huangruiteng/loopx '''loopx''' ] - Noyau d'ingénierie de boucle légère pour équipes d'agents IA à long terme. Agnostique aux agents de codage, avec objectifs durables et transferts vérifiables.&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] : Transforme n importe quel PDF technique en skill Claude Code - prêt à étudier et référencer.&lt;br /&gt;
* [https://github.com/firecrawl/pdf-inspector '''pdf-inspector'''] : Bibliothèque Rust ultra-rapide pour l'inspection, la classification et l'extraction de texte PDF — detection intelligente scans vs texte pour le routing RAG. (8 600+ ★)&lt;br /&gt;
* [https://github.com/vitali87/code-graph-rag '''Code-Graph-RAG'''] : RAG pour monorepo via graphe de connaissances — interrogez, comprenez et editez votre codebase multi-langage avec l'IA et Tree-sitter. MCP server compatible.&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/antirez/ds4 '''ds4'''] : Moteur inference local DeepSeek 4 Flash/PRO pour Metal, CUDA et ROCm - par Antirez.&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D '''GROQ LLM accelerator''']&lt;br /&gt;
* [https://github.com/Comfy-Org/ComfyUI '''ComfyUI''' ] : Interface nodes pour modèles de diffusion — API, backend modular et exécuteur de workflows locaux pour génération d'images et IA visuelle. (125 000+ ★)&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
* [https://github.com/agavra/tuicr '''tuicr''' ] - Revue de code TUI avec raccourcis vim.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/goauthentik/authentik '''authentik'''] : Fournisseur d'identite open-source auto-herberge — SSO, MFA, SCIM, federation d'identite. L'alternative open-source a Okta/Auth0. (54k+ ★)&lt;br /&gt;
* [https://github.com/semantica-agi/semantica '''semantica''' ] - Infrastructure graph-native pour contexte et systèmes IA responsables.&lt;br /&gt;
* [https://github.com/drawdb-io/drawdb '''drawdb''' ] - Éditeur de diagrammes de base de données en ligne gratuit et intuitif.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2094</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2094"/>
		<updated>2026-08-11T01:27:12Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Ajout projets trending GitHub - Semaine 08/11/2026: loopx (Agents), semantica &amp;amp; drawdb (Infrastructure)&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher&lt;br /&gt;
* [https://github.com/different-ai/openwork '''openwork''' ] : Alternative open-source à Claude Cowork — app desktop de partage de workflows IA avec support MCP multi-agents, gestion d'équipes et marketplace de skills. (21 600 ★)&lt;br /&gt;
* [https://github.com/google/skills '''Agent Skills Google''' ] : Collection officielle de skills IA pour produits et technologies Google Cloud — intégration Claude Code, Codex, Antigravity CLI et agents compatibles. (16 742 ★)&lt;br /&gt;
 des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com/docs '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 02/06 - 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] : API de traduction multilingue temps réel — PR #3 fix timeout (gestion réponses lentes), PR #2 URLs configurables via .env.&lt;br /&gt;
* [https://github.com/ynotopec/GLM-4.7-Flash-vLLM '''GLM-4.7-Flash-vLLM'''] : Déploiement vLLM GLM-4.7 Flash.&lt;br /&gt;
* [https://github.com/ynotopec/qwen3.5-fp8-vllm '''qwen3.5-fp8-vllm'''] : Déploiement Qwen 3.5 FP8 via vLLM.&lt;br /&gt;
* [https://github.com/ynotopec/api-rag '''api-rag'''] : API RAG (recherche sémantique) — fix support FAISS SVE, suppression logs verbeux, wait-for-APIS avant build index vectoriel.&lt;br /&gt;
* [https://github.com/ynotopec/api-diarization '''api-diarization'''] : API diarisation audio — install améliorée, ergonomics auth, documentation setup projet.&lt;br /&gt;
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] : API FastAPI ASR ( Whisper ) — refonte complète, installer uv, docs runtime.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-openmed '''privacy-proxy-openmed'''] : Privacy proxy OpenMed-based — LLM upstream optionnel (mode privacy-only).&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''privacy-proxy-gliner2'''] : Privacy proxy GLiNER2-based — détection entités nommées, upstream LLM optionnel, PR #1 en cours.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] : Proxy LLM custom — paramètres modèle configurables, alias ai-default, gestion erreurs upstream.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Priorités &amp;amp; Veille ==&lt;br /&gt;
&lt;br /&gt;
Les priorités ci-dessous sont les seules entrées de veille conservées. Chaque item doit avoir un **statu** : *en cours*, *testé*, *retenu*, *abandonné*. Les projets listés sans action réelle sont retirés à chaque révision.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration d'inférence multi-nœuds.&lt;br /&gt;
* [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning et benchmark réaliste.&lt;br /&gt;
* [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques pour les LLM.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix / token en temps réel.&lt;br /&gt;
* '''Qualité auto''' : précision / taux d'hallucination sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] : routage sémantique de requêtes vers les bons modèles.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache'''] : couche KV cache pour réduire la latence d'inférence.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] : orchestration de workflows critiques avec persistence.&lt;br /&gt;
* [https://github.com/lyogavin/airllm '''airLLM'''] : inférence LLM sur GPU 4 Go — pour réduire les coûts de inférence locale.&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute'''] : passerelle multi-providers (50+ gratuits), compression 15-95% de tokens — à évaluer comme fallback.&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Assistants IA &amp;amp; Agents ===&lt;br /&gt;
* [https://github.com/TencentCloud/TencentDB-Agent-Memory '''TencentDB-Agent-Memory'''] : Hub de mémoire d équipe pour agents IA - transforme conversations, docs et code en assets réutilisables.&lt;br /&gt;
* [https://github.com/unclebob/swarm-forge '''Swarm Forge'''] : Outil simple de coordination de plusieurs agents IA.&lt;br /&gt;
* [https://github.com/embabel/embabel-agent '''Embabel Agent'''] : Framework agent pour JVM (Kotlin/Java).&lt;br /&gt;
* [https://github.com/zhaoxuya520/reverse-skill '''reverse-skill'''] : Skill Router Pack pour reverse engineering et pentesting - routage IA, bootstrapping toolchain auto.&lt;br /&gt;
* [https://github.com/esengine/DeepSeek-Reasonix '''DeepSeek-Reasonix'''] : Agent de codage terminal natif DeepSeek, optimisé pour la stabilité du prefix-cache.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd'''] : Skill qui empêche l agent de coder de cacher la réponse - output clair et orienté action.&lt;br /&gt;
* [https://github.com/block/buzz '''Buzz (Block)'''] : Plateforme de communication type hive mind pour collaboration multi-agents.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/livekit/agents '''LiveKit Agents'''] : Framework Python/Node.js pour agents IA vocaux en temps réel — build, deployer et scaler des agents multimodaux en production. (1 100+ ★)&lt;br /&gt;
* [https://github.com/huangruiteng/loopx '''loopx''' ] - Noyau d'ingénierie de boucle légère pour équipes d'agents IA à long terme. Agnostique aux agents de codage, avec objectifs durables et transferts vérifiables.&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] : Transforme n importe quel PDF technique en skill Claude Code - prêt à étudier et référencer.&lt;br /&gt;
* [https://github.com/firecrawl/pdf-inspector '''pdf-inspector'''] : Bibliothèque Rust ultra-rapide pour l'inspection, la classification et l'extraction de texte PDF — detection intelligente scans vs texte pour le routing RAG. (8 600+ ★)&lt;br /&gt;
* [https://github.com/vitali87/code-graph-rag '''Code-Graph-RAG'''] : RAG pour monorepo via graphe de connaissances — interrogez, comprenez et editez votre codebase multi-langage avec l'IA et Tree-sitter. MCP server compatible.&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/antirez/ds4 '''ds4'''] : Moteur inference local DeepSeek 4 Flash/PRO pour Metal, CUDA et ROCm - par Antirez.&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D '''GROQ LLM accelerator''']&lt;br /&gt;
* [https://github.com/Comfy-Org/ComfyUI '''ComfyUI''' ] : Interface nodes pour modèles de diffusion — API, backend modular et exécuteur de workflows locaux pour génération d'images et IA visuelle. (125 000+ ★)&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
* [https://github.com/agavra/tuicr '''tuicr''' ] - Revue de code TUI avec raccourcis vim.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/goauthentik/authentik '''authentik'''] : Fournisseur d'identite open-source auto-herberge — SSO, MFA, SCIM, federation d'identite. L'alternative open-source a Okta/Auth0. (54k+ ★)&lt;br /&gt;
* [https://github.com/semantica-agi/semantica '''semantica''' ] - Infrastructure graph-native pour contexte et systèmes IA responsables.&lt;br /&gt;
* [https://github.com/drawdb-io/drawdb '''drawdb''' ] - Éditeur de diagrammes de base de données en ligne gratuit et intuitif.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2093</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2093"/>
		<updated>2026-08-10T11:02:41Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: /* Matériel IA &amp;amp; GPU */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher&lt;br /&gt;
* [https://github.com/different-ai/openwork '''openwork''' ] : Alternative open-source à Claude Cowork — app desktop de partage de workflows IA avec support MCP multi-agents, gestion d'équipes et marketplace de skills. (21 600 ★)&lt;br /&gt;
* [https://github.com/google/skills '''Agent Skills Google''' ] : Collection officielle de skills IA pour produits et technologies Google Cloud — intégration Claude Code, Codex, Antigravity CLI et agents compatibles. (16 742 ★)&lt;br /&gt;
 des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com/docs '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 02/06 - 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] : API de traduction multilingue temps réel — PR #3 fix timeout (gestion réponses lentes), PR #2 URLs configurables via .env.&lt;br /&gt;
* [https://github.com/ynotopec/GLM-4.7-Flash-vLLM '''GLM-4.7-Flash-vLLM'''] : Déploiement vLLM GLM-4.7 Flash.&lt;br /&gt;
* [https://github.com/ynotopec/qwen3.5-fp8-vllm '''qwen3.5-fp8-vllm'''] : Déploiement Qwen 3.5 FP8 via vLLM.&lt;br /&gt;
* [https://github.com/ynotopec/api-rag '''api-rag'''] : API RAG (recherche sémantique) — fix support FAISS SVE, suppression logs verbeux, wait-for-APIS avant build index vectoriel.&lt;br /&gt;
* [https://github.com/ynotopec/api-diarization '''api-diarization'''] : API diarisation audio — install améliorée, ergonomics auth, documentation setup projet.&lt;br /&gt;
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] : API FastAPI ASR ( Whisper ) — refonte complète, installer uv, docs runtime.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-openmed '''privacy-proxy-openmed'''] : Privacy proxy OpenMed-based — LLM upstream optionnel (mode privacy-only).&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''privacy-proxy-gliner2'''] : Privacy proxy GLiNER2-based — détection entités nommées, upstream LLM optionnel, PR #1 en cours.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] : Proxy LLM custom — paramètres modèle configurables, alias ai-default, gestion erreurs upstream.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Priorités &amp;amp; Veille ==&lt;br /&gt;
&lt;br /&gt;
Les priorités ci-dessous sont les seules entrées de veille conservées. Chaque item doit avoir un **statu** : *en cours*, *testé*, *retenu*, *abandonné*. Les projets listés sans action réelle sont retirés à chaque révision.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration d'inférence multi-nœuds.&lt;br /&gt;
* [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning et benchmark réaliste.&lt;br /&gt;
* [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques pour les LLM.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix / token en temps réel.&lt;br /&gt;
* '''Qualité auto''' : précision / taux d'hallucination sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] : routage sémantique de requêtes vers les bons modèles.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache'''] : couche KV cache pour réduire la latence d'inférence.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] : orchestration de workflows critiques avec persistence.&lt;br /&gt;
* [https://github.com/lyogavin/airllm '''airLLM'''] : inférence LLM sur GPU 4 Go — pour réduire les coûts de inférence locale.&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute'''] : passerelle multi-providers (50+ gratuits), compression 15-95% de tokens — à évaluer comme fallback.&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Assistants IA &amp;amp; Agents ===&lt;br /&gt;
* [https://github.com/TencentCloud/TencentDB-Agent-Memory '''TencentDB-Agent-Memory'''] : Hub de mémoire d équipe pour agents IA - transforme conversations, docs et code en assets réutilisables.&lt;br /&gt;
* [https://github.com/unclebob/swarm-forge '''Swarm Forge'''] : Outil simple de coordination de plusieurs agents IA.&lt;br /&gt;
* [https://github.com/embabel/embabel-agent '''Embabel Agent'''] : Framework agent pour JVM (Kotlin/Java).&lt;br /&gt;
* [https://github.com/zhaoxuya520/reverse-skill '''reverse-skill'''] : Skill Router Pack pour reverse engineering et pentesting - routage IA, bootstrapping toolchain auto.&lt;br /&gt;
* [https://github.com/esengine/DeepSeek-Reasonix '''DeepSeek-Reasonix'''] : Agent de codage terminal natif DeepSeek, optimisé pour la stabilité du prefix-cache.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd'''] : Skill qui empêche l agent de coder de cacher la réponse - output clair et orienté action.&lt;br /&gt;
* [https://github.com/block/buzz '''Buzz (Block)'''] : Plateforme de communication type hive mind pour collaboration multi-agents.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/livekit/agents '''LiveKit Agents'''] : Framework Python/Node.js pour agents IA vocaux en temps réel — build, deployer et scaler des agents multimodaux en production. (1 100+ ★)&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] : Transforme n importe quel PDF technique en skill Claude Code - prêt à étudier et référencer.&lt;br /&gt;
* [https://github.com/firecrawl/pdf-inspector '''pdf-inspector'''] : Bibliothèque Rust ultra-rapide pour l'inspection, la classification et l'extraction de texte PDF — detection intelligente scans vs texte pour le routing RAG. (8 600+ ★)&lt;br /&gt;
* [https://github.com/vitali87/code-graph-rag '''Code-Graph-RAG'''] : RAG pour monorepo via graphe de connaissances — interrogez, comprenez et editez votre codebase multi-langage avec l'IA et Tree-sitter. MCP server compatible.&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/antirez/ds4 '''ds4'''] : Moteur inference local DeepSeek 4 Flash/PRO pour Metal, CUDA et ROCm - par Antirez.&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D '''GROQ LLM accelerator''']&lt;br /&gt;
* [https://github.com/Comfy-Org/ComfyUI '''ComfyUI''' ] : Interface nodes pour modèles de diffusion — API, backend modular et exécuteur de workflows locaux pour génération d'images et IA visuelle. (125 000+ ★)&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
* [https://github.com/agavra/tuicr '''tuicr''' ] - Revue de code TUI avec raccourcis vim.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/goauthentik/authentik '''authentik'''] : Fournisseur d'identite open-source auto-herberge — SSO, MFA, SCIM, federation d'identite. L'alternative open-source a Okta/Auth0. (54k+ ★)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2092</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2092"/>
		<updated>2026-08-10T11:02:26Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: /* Matériel IA &amp;amp; GPU */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher&lt;br /&gt;
* [https://github.com/different-ai/openwork '''openwork''' ] : Alternative open-source à Claude Cowork — app desktop de partage de workflows IA avec support MCP multi-agents, gestion d'équipes et marketplace de skills. (21 600 ★)&lt;br /&gt;
* [https://github.com/google/skills '''Agent Skills Google''' ] : Collection officielle de skills IA pour produits et technologies Google Cloud — intégration Claude Code, Codex, Antigravity CLI et agents compatibles. (16 742 ★)&lt;br /&gt;
 des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com/docs '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 02/06 - 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] : API de traduction multilingue temps réel — PR #3 fix timeout (gestion réponses lentes), PR #2 URLs configurables via .env.&lt;br /&gt;
* [https://github.com/ynotopec/GLM-4.7-Flash-vLLM '''GLM-4.7-Flash-vLLM'''] : Déploiement vLLM GLM-4.7 Flash.&lt;br /&gt;
* [https://github.com/ynotopec/qwen3.5-fp8-vllm '''qwen3.5-fp8-vllm'''] : Déploiement Qwen 3.5 FP8 via vLLM.&lt;br /&gt;
* [https://github.com/ynotopec/api-rag '''api-rag'''] : API RAG (recherche sémantique) — fix support FAISS SVE, suppression logs verbeux, wait-for-APIS avant build index vectoriel.&lt;br /&gt;
* [https://github.com/ynotopec/api-diarization '''api-diarization'''] : API diarisation audio — install améliorée, ergonomics auth, documentation setup projet.&lt;br /&gt;
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] : API FastAPI ASR ( Whisper ) — refonte complète, installer uv, docs runtime.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-openmed '''privacy-proxy-openmed'''] : Privacy proxy OpenMed-based — LLM upstream optionnel (mode privacy-only).&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''privacy-proxy-gliner2'''] : Privacy proxy GLiNER2-based — détection entités nommées, upstream LLM optionnel, PR #1 en cours.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] : Proxy LLM custom — paramètres modèle configurables, alias ai-default, gestion erreurs upstream.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Priorités &amp;amp; Veille ==&lt;br /&gt;
&lt;br /&gt;
Les priorités ci-dessous sont les seules entrées de veille conservées. Chaque item doit avoir un **statu** : *en cours*, *testé*, *retenu*, *abandonné*. Les projets listés sans action réelle sont retirés à chaque révision.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration d'inférence multi-nœuds.&lt;br /&gt;
* [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning et benchmark réaliste.&lt;br /&gt;
* [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques pour les LLM.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix / token en temps réel.&lt;br /&gt;
* '''Qualité auto''' : précision / taux d'hallucination sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] : routage sémantique de requêtes vers les bons modèles.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache'''] : couche KV cache pour réduire la latence d'inférence.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] : orchestration de workflows critiques avec persistence.&lt;br /&gt;
* [https://github.com/lyogavin/airllm '''airLLM'''] : inférence LLM sur GPU 4 Go — pour réduire les coûts de inférence locale.&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute'''] : passerelle multi-providers (50+ gratuits), compression 15-95% de tokens — à évaluer comme fallback.&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Assistants IA &amp;amp; Agents ===&lt;br /&gt;
* [https://github.com/TencentCloud/TencentDB-Agent-Memory '''TencentDB-Agent-Memory'''] : Hub de mémoire d équipe pour agents IA - transforme conversations, docs et code en assets réutilisables.&lt;br /&gt;
* [https://github.com/unclebob/swarm-forge '''Swarm Forge'''] : Outil simple de coordination de plusieurs agents IA.&lt;br /&gt;
* [https://github.com/embabel/embabel-agent '''Embabel Agent'''] : Framework agent pour JVM (Kotlin/Java).&lt;br /&gt;
* [https://github.com/zhaoxuya520/reverse-skill '''reverse-skill'''] : Skill Router Pack pour reverse engineering et pentesting - routage IA, bootstrapping toolchain auto.&lt;br /&gt;
* [https://github.com/esengine/DeepSeek-Reasonix '''DeepSeek-Reasonix'''] : Agent de codage terminal natif DeepSeek, optimisé pour la stabilité du prefix-cache.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd'''] : Skill qui empêche l agent de coder de cacher la réponse - output clair et orienté action.&lt;br /&gt;
* [https://github.com/block/buzz '''Buzz (Block)'''] : Plateforme de communication type hive mind pour collaboration multi-agents.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/livekit/agents '''LiveKit Agents'''] : Framework Python/Node.js pour agents IA vocaux en temps réel — build, deployer et scaler des agents multimodaux en production. (1 100+ ★)&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] : Transforme n importe quel PDF technique en skill Claude Code - prêt à étudier et référencer.&lt;br /&gt;
* [https://github.com/firecrawl/pdf-inspector '''pdf-inspector'''] : Bibliothèque Rust ultra-rapide pour l'inspection, la classification et l'extraction de texte PDF — detection intelligente scans vs texte pour le routing RAG. (8 600+ ★)&lt;br /&gt;
* [https://github.com/vitali87/code-graph-rag '''Code-Graph-RAG'''] : RAG pour monorepo via graphe de connaissances — interrogez, comprenez et editez votre codebase multi-langage avec l'IA et Tree-sitter. MCP server compatible.&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/antirez/ds4 '''ds4'''] : Moteur inference local DeepSeek 4 Flash/PRO pour Metal, CUDA et ROCm - par Antirez.&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
* [https://github.com/Comfy-Org/ComfyUI '''ComfyUI''' ] : Interface nodes pour modèles de diffusion — API, backend modular et exécuteur de workflows locaux pour génération d'images et IA visuelle. (125 000+ ★)&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
* [https://github.com/agavra/tuicr '''tuicr''' ] - Revue de code TUI avec raccourcis vim.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/goauthentik/authentik '''authentik'''] : Fournisseur d'identite open-source auto-herberge — SSO, MFA, SCIM, federation d'identite. L'alternative open-source a Okta/Auth0. (54k+ ★)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2091</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2091"/>
		<updated>2026-08-10T01:28:44Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Correction placement sections: pdf-inspector et Code-Graph-RAG vers RAG &amp;amp; Traitement de Documents&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher&lt;br /&gt;
* [https://github.com/different-ai/openwork '''openwork''' ] : Alternative open-source à Claude Cowork — app desktop de partage de workflows IA avec support MCP multi-agents, gestion d'équipes et marketplace de skills. (21 600 ★)&lt;br /&gt;
* [https://github.com/google/skills '''Agent Skills Google''' ] : Collection officielle de skills IA pour produits et technologies Google Cloud — intégration Claude Code, Codex, Antigravity CLI et agents compatibles. (16 742 ★)&lt;br /&gt;
 des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com/docs '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 02/06 - 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] : API de traduction multilingue temps réel — PR #3 fix timeout (gestion réponses lentes), PR #2 URLs configurables via .env.&lt;br /&gt;
* [https://github.com/ynotopec/GLM-4.7-Flash-vLLM '''GLM-4.7-Flash-vLLM'''] : Déploiement vLLM GLM-4.7 Flash.&lt;br /&gt;
* [https://github.com/ynotopec/qwen3.5-fp8-vllm '''qwen3.5-fp8-vllm'''] : Déploiement Qwen 3.5 FP8 via vLLM.&lt;br /&gt;
* [https://github.com/ynotopec/api-rag '''api-rag'''] : API RAG (recherche sémantique) — fix support FAISS SVE, suppression logs verbeux, wait-for-APIS avant build index vectoriel.&lt;br /&gt;
* [https://github.com/ynotopec/api-diarization '''api-diarization'''] : API diarisation audio — install améliorée, ergonomics auth, documentation setup projet.&lt;br /&gt;
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] : API FastAPI ASR ( Whisper ) — refonte complète, installer uv, docs runtime.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-openmed '''privacy-proxy-openmed'''] : Privacy proxy OpenMed-based — LLM upstream optionnel (mode privacy-only).&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''privacy-proxy-gliner2'''] : Privacy proxy GLiNER2-based — détection entités nommées, upstream LLM optionnel, PR #1 en cours.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] : Proxy LLM custom — paramètres modèle configurables, alias ai-default, gestion erreurs upstream.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Priorités &amp;amp; Veille ==&lt;br /&gt;
&lt;br /&gt;
Les priorités ci-dessous sont les seules entrées de veille conservées. Chaque item doit avoir un **statu** : *en cours*, *testé*, *retenu*, *abandonné*. Les projets listés sans action réelle sont retirés à chaque révision.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration d'inférence multi-nœuds.&lt;br /&gt;
* [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning et benchmark réaliste.&lt;br /&gt;
* [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques pour les LLM.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix / token en temps réel.&lt;br /&gt;
* '''Qualité auto''' : précision / taux d'hallucination sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] : routage sémantique de requêtes vers les bons modèles.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache'''] : couche KV cache pour réduire la latence d'inférence.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] : orchestration de workflows critiques avec persistence.&lt;br /&gt;
* [https://github.com/lyogavin/airllm '''airLLM'''] : inférence LLM sur GPU 4 Go — pour réduire les coûts de inférence locale.&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute'''] : passerelle multi-providers (50+ gratuits), compression 15-95% de tokens — à évaluer comme fallback.&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Assistants IA &amp;amp; Agents ===&lt;br /&gt;
* [https://github.com/TencentCloud/TencentDB-Agent-Memory '''TencentDB-Agent-Memory'''] : Hub de mémoire d équipe pour agents IA - transforme conversations, docs et code en assets réutilisables.&lt;br /&gt;
* [https://github.com/unclebob/swarm-forge '''Swarm Forge'''] : Outil simple de coordination de plusieurs agents IA.&lt;br /&gt;
* [https://github.com/embabel/embabel-agent '''Embabel Agent'''] : Framework agent pour JVM (Kotlin/Java).&lt;br /&gt;
* [https://github.com/zhaoxuya520/reverse-skill '''reverse-skill'''] : Skill Router Pack pour reverse engineering et pentesting - routage IA, bootstrapping toolchain auto.&lt;br /&gt;
* [https://github.com/esengine/DeepSeek-Reasonix '''DeepSeek-Reasonix'''] : Agent de codage terminal natif DeepSeek, optimisé pour la stabilité du prefix-cache.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd'''] : Skill qui empêche l agent de coder de cacher la réponse - output clair et orienté action.&lt;br /&gt;
* [https://github.com/block/buzz '''Buzz (Block)'''] : Plateforme de communication type hive mind pour collaboration multi-agents.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/livekit/agents '''LiveKit Agents'''] : Framework Python/Node.js pour agents IA vocaux en temps réel — build, deployer et scaler des agents multimodaux en production. (1 100+ ★)&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] : Transforme n importe quel PDF technique en skill Claude Code - prêt à étudier et référencer.&lt;br /&gt;
* [https://github.com/firecrawl/pdf-inspector '''pdf-inspector'''] : Bibliothèque Rust ultra-rapide pour l'inspection, la classification et l'extraction de texte PDF — detection intelligente scans vs texte pour le routing RAG. (8 600+ ★)&lt;br /&gt;
* [https://github.com/vitali87/code-graph-rag '''Code-Graph-RAG'''] : RAG pour monorepo via graphe de connaissances — interrogez, comprenez et editez votre codebase multi-langage avec l'IA et Tree-sitter. MCP server compatible.&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/antirez/ds4 '''ds4'''] : Moteur inference local DeepSeek 4 Flash/PRO pour Metal, CUDA et ROCm - par Antirez.&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQm&lt;br /&gt;
* [https://github.com/Comfy-Org/ComfyUI '''ComfyUI''' ] : Interface nodes pour modèles de diffusion — API, backend modular et exécuteur de workflows locaux pour génération d'images et IA visuelle. (125 000+ ★)&lt;br /&gt;
Fw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
* [https://github.com/agavra/tuicr '''tuicr''' ] - Revue de code TUI avec raccourcis vim.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/goauthentik/authentik '''authentik'''] : Fournisseur d'identite open-source auto-herberge — SSO, MFA, SCIM, federation d'identite. L'alternative open-source a Okta/Auth0. (54k+ ★)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2090</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2090"/>
		<updated>2026-08-10T01:25:02Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Ajout hebdomadaire GitHub Trending: LiveKit Agents, pdf-inspector, Code-Graph-RAG, authentik&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher&lt;br /&gt;
* [https://github.com/different-ai/openwork '''openwork''' ] : Alternative open-source à Claude Cowork — app desktop de partage de workflows IA avec support MCP multi-agents, gestion d'équipes et marketplace de skills. (21 600 ★)&lt;br /&gt;
* [https://github.com/google/skills '''Agent Skills Google''' ] : Collection officielle de skills IA pour produits et technologies Google Cloud — intégration Claude Code, Codex, Antigravity CLI et agents compatibles. (16 742 ★)&lt;br /&gt;
 des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com/docs '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 02/06 - 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] : API de traduction multilingue temps réel — PR #3 fix timeout (gestion réponses lentes), PR #2 URLs configurables via .env.&lt;br /&gt;
* [https://github.com/ynotopec/GLM-4.7-Flash-vLLM '''GLM-4.7-Flash-vLLM'''] : Déploiement vLLM GLM-4.7 Flash.&lt;br /&gt;
* [https://github.com/ynotopec/qwen3.5-fp8-vllm '''qwen3.5-fp8-vllm'''] : Déploiement Qwen 3.5 FP8 via vLLM.&lt;br /&gt;
* [https://github.com/ynotopec/api-rag '''api-rag'''] : API RAG (recherche sémantique) — fix support FAISS SVE, suppression logs verbeux, wait-for-APIS avant build index vectoriel.&lt;br /&gt;
* [https://github.com/ynotopec/api-diarization '''api-diarization'''] : API diarisation audio — install améliorée, ergonomics auth, documentation setup projet.&lt;br /&gt;
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] : API FastAPI ASR ( Whisper ) — refonte complète, installer uv, docs runtime.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-openmed '''privacy-proxy-openmed'''] : Privacy proxy OpenMed-based — LLM upstream optionnel (mode privacy-only).&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''privacy-proxy-gliner2'''] : Privacy proxy GLiNER2-based — détection entités nommées, upstream LLM optionnel, PR #1 en cours.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] : Proxy LLM custom — paramètres modèle configurables, alias ai-default, gestion erreurs upstream.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Priorités &amp;amp; Veille ==&lt;br /&gt;
&lt;br /&gt;
Les priorités ci-dessous sont les seules entrées de veille conservées. Chaque item doit avoir un **statu** : *en cours*, *testé*, *retenu*, *abandonné*. Les projets listés sans action réelle sont retirés à chaque révision.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration d'inférence multi-nœuds.&lt;br /&gt;
* [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning et benchmark réaliste.&lt;br /&gt;
* [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques pour les LLM.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix / token en temps réel.&lt;br /&gt;
* '''Qualité auto''' : précision / taux d'hallucination sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] : routage sémantique de requêtes vers les bons modèles.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache'''] : couche KV cache pour réduire la latence d'inférence.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] : orchestration de workflows critiques avec persistence.&lt;br /&gt;
* [https://github.com/lyogavin/airllm '''airLLM'''] : inférence LLM sur GPU 4 Go — pour réduire les coûts de inférence locale.&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute'''] : passerelle multi-providers (50+ gratuits), compression 15-95% de tokens — à évaluer comme fallback.&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Assistants IA &amp;amp; Agents ===&lt;br /&gt;
* [https://github.com/TencentCloud/TencentDB-Agent-Memory '''TencentDB-Agent-Memory'''] : Hub de mémoire d équipe pour agents IA - transforme conversations, docs et code en assets réutilisables.&lt;br /&gt;
* [https://github.com/unclebob/swarm-forge '''Swarm Forge'''] : Outil simple de coordination de plusieurs agents IA.&lt;br /&gt;
* [https://github.com/embabel/embabel-agent '''Embabel Agent'''] : Framework agent pour JVM (Kotlin/Java).&lt;br /&gt;
* [https://github.com/zhaoxuya520/reverse-skill '''reverse-skill'''] : Skill Router Pack pour reverse engineering et pentesting - routage IA, bootstrapping toolchain auto.&lt;br /&gt;
* [https://github.com/esengine/DeepSeek-Reasonix '''DeepSeek-Reasonix'''] : Agent de codage terminal natif DeepSeek, optimisé pour la stabilité du prefix-cache.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd'''] : Skill qui empêche l agent de coder de cacher la réponse - output clair et orienté action.&lt;br /&gt;
* [https://github.com/block/buzz '''Buzz (Block)'''] : Plateforme de communication type hive mind pour collaboration multi-agents.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/livekit/agents '''LiveKit Agents'''] : Framework Python/Node.js pour agents IA vocaux en temps réel — build, deployer et scaler des agents multimodaux en production. (1 100+ ★)&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/firecrawl/pdf-inspector '''pdf-inspector'''] : Bibliothèque Rust ultra-rapide pour l'inspection, la classification et l'extraction de texte PDF — detection intelligente scans vs texte pour le routing RAG. (8 600+ ★)&lt;br /&gt;
* [https://github.com/vitali87/code-graph-rag '''Code-Graph-RAG'''] : RAG pour monorepo via graphe de connaissances — interrogez, comprenez et editez votre codebase multi-langage avec l'IA et Tree-sitter. MCP server compatible.&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] : Transforme n importe quel PDF technique en skill Claude Code - prêt à étudier et référencer.&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/antirez/ds4 '''ds4'''] : Moteur inference local DeepSeek 4 Flash/PRO pour Metal, CUDA et ROCm - par Antirez.&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQm&lt;br /&gt;
* [https://github.com/Comfy-Org/ComfyUI '''ComfyUI''' ] : Interface nodes pour modèles de diffusion — API, backend modular et exécuteur de workflows locaux pour génération d'images et IA visuelle. (125 000+ ★)&lt;br /&gt;
Fw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
* [https://github.com/agavra/tuicr '''tuicr''' ] - Revue de code TUI avec raccourcis vim.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/goauthentik/authentik '''authentik'''] : Fournisseur d'identite open-source auto-herberge — SSO, MFA, SCIM, federation d'identite. L'alternative open-source a Okta/Auth0. (54k+ ★)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2089</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2089"/>
		<updated>2026-08-09T01:21:31Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Ajout projets trending GitHub semanal : openwork, Agent Skills Google, ComfyUI&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher&lt;br /&gt;
* [https://github.com/different-ai/openwork '''openwork''' ] : Alternative open-source à Claude Cowork — app desktop de partage de workflows IA avec support MCP multi-agents, gestion d'équipes et marketplace de skills. (21 600 ★)&lt;br /&gt;
* [https://github.com/google/skills '''Agent Skills Google''' ] : Collection officielle de skills IA pour produits et technologies Google Cloud — intégration Claude Code, Codex, Antigravity CLI et agents compatibles. (16 742 ★)&lt;br /&gt;
 des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com/docs '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 02/06 - 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] : API de traduction multilingue temps réel — PR #3 fix timeout (gestion réponses lentes), PR #2 URLs configurables via .env.&lt;br /&gt;
* [https://github.com/ynotopec/GLM-4.7-Flash-vLLM '''GLM-4.7-Flash-vLLM'''] : Déploiement vLLM GLM-4.7 Flash.&lt;br /&gt;
* [https://github.com/ynotopec/qwen3.5-fp8-vllm '''qwen3.5-fp8-vllm'''] : Déploiement Qwen 3.5 FP8 via vLLM.&lt;br /&gt;
* [https://github.com/ynotopec/api-rag '''api-rag'''] : API RAG (recherche sémantique) — fix support FAISS SVE, suppression logs verbeux, wait-for-APIS avant build index vectoriel.&lt;br /&gt;
* [https://github.com/ynotopec/api-diarization '''api-diarization'''] : API diarisation audio — install améliorée, ergonomics auth, documentation setup projet.&lt;br /&gt;
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] : API FastAPI ASR ( Whisper ) — refonte complète, installer uv, docs runtime.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-openmed '''privacy-proxy-openmed'''] : Privacy proxy OpenMed-based — LLM upstream optionnel (mode privacy-only).&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''privacy-proxy-gliner2'''] : Privacy proxy GLiNER2-based — détection entités nommées, upstream LLM optionnel, PR #1 en cours.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] : Proxy LLM custom — paramètres modèle configurables, alias ai-default, gestion erreurs upstream.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Priorités &amp;amp; Veille ==&lt;br /&gt;
&lt;br /&gt;
Les priorités ci-dessous sont les seules entrées de veille conservées. Chaque item doit avoir un **statu** : *en cours*, *testé*, *retenu*, *abandonné*. Les projets listés sans action réelle sont retirés à chaque révision.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration d'inférence multi-nœuds.&lt;br /&gt;
* [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning et benchmark réaliste.&lt;br /&gt;
* [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques pour les LLM.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix / token en temps réel.&lt;br /&gt;
* '''Qualité auto''' : précision / taux d'hallucination sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] : routage sémantique de requêtes vers les bons modèles.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache'''] : couche KV cache pour réduire la latence d'inférence.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] : orchestration de workflows critiques avec persistence.&lt;br /&gt;
* [https://github.com/lyogavin/airllm '''airLLM'''] : inférence LLM sur GPU 4 Go — pour réduire les coûts de inférence locale.&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute'''] : passerelle multi-providers (50+ gratuits), compression 15-95% de tokens — à évaluer comme fallback.&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Assistants IA &amp;amp; Agents ===&lt;br /&gt;
* [https://github.com/TencentCloud/TencentDB-Agent-Memory '''TencentDB-Agent-Memory'''] : Hub de mémoire d équipe pour agents IA - transforme conversations, docs et code en assets réutilisables.&lt;br /&gt;
* [https://github.com/unclebob/swarm-forge '''Swarm Forge'''] : Outil simple de coordination de plusieurs agents IA.&lt;br /&gt;
* [https://github.com/embabel/embabel-agent '''Embabel Agent'''] : Framework agent pour JVM (Kotlin/Java).&lt;br /&gt;
* [https://github.com/zhaoxuya520/reverse-skill '''reverse-skill'''] : Skill Router Pack pour reverse engineering et pentesting - routage IA, bootstrapping toolchain auto.&lt;br /&gt;
* [https://github.com/esengine/DeepSeek-Reasonix '''DeepSeek-Reasonix'''] : Agent de codage terminal natif DeepSeek, optimisé pour la stabilité du prefix-cache.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd'''] : Skill qui empêche l agent de coder de cacher la réponse - output clair et orienté action.&lt;br /&gt;
* [https://github.com/block/buzz '''Buzz (Block)'''] : Plateforme de communication type hive mind pour collaboration multi-agents.&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] : Transforme n importe quel PDF technique en skill Claude Code - prêt à étudier et référencer.&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/antirez/ds4 '''ds4'''] : Moteur inference local DeepSeek 4 Flash/PRO pour Metal, CUDA et ROCm - par Antirez.&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQm&lt;br /&gt;
* [https://github.com/Comfy-Org/ComfyUI '''ComfyUI''' ] : Interface nodes pour modèles de diffusion — API, backend modular et exécuteur de workflows locaux pour génération d'images et IA visuelle. (125 000+ ★)&lt;br /&gt;
Fw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
* [https://github.com/agavra/tuicr '''tuicr''' ] - Revue de code TUI avec raccourcis vim.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2088</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2088"/>
		<updated>2026-08-08T01:23:28Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Veille hebdo GitHub: ajout de 9 projets trending dans les sections catégorisées (Agents IA, RAG, Infrastructure)&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com/docs '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 02/06 - 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] : API de traduction multilingue temps réel — PR #3 fix timeout (gestion réponses lentes), PR #2 URLs configurables via .env.&lt;br /&gt;
* [https://github.com/ynotopec/GLM-4.7-Flash-vLLM '''GLM-4.7-Flash-vLLM'''] : Déploiement vLLM GLM-4.7 Flash.&lt;br /&gt;
* [https://github.com/ynotopec/qwen3.5-fp8-vllm '''qwen3.5-fp8-vllm'''] : Déploiement Qwen 3.5 FP8 via vLLM.&lt;br /&gt;
* [https://github.com/ynotopec/api-rag '''api-rag'''] : API RAG (recherche sémantique) — fix support FAISS SVE, suppression logs verbeux, wait-for-APIS avant build index vectoriel.&lt;br /&gt;
* [https://github.com/ynotopec/api-diarization '''api-diarization'''] : API diarisation audio — install améliorée, ergonomics auth, documentation setup projet.&lt;br /&gt;
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] : API FastAPI ASR ( Whisper ) — refonte complète, installer uv, docs runtime.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-openmed '''privacy-proxy-openmed'''] : Privacy proxy OpenMed-based — LLM upstream optionnel (mode privacy-only).&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''privacy-proxy-gliner2'''] : Privacy proxy GLiNER2-based — détection entités nommées, upstream LLM optionnel, PR #1 en cours.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] : Proxy LLM custom — paramètres modèle configurables, alias ai-default, gestion erreurs upstream.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Priorités &amp;amp; Veille ==&lt;br /&gt;
&lt;br /&gt;
Les priorités ci-dessous sont les seules entrées de veille conservées. Chaque item doit avoir un **statu** : *en cours*, *testé*, *retenu*, *abandonné*. Les projets listés sans action réelle sont retirés à chaque révision.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration d'inférence multi-nœuds.&lt;br /&gt;
* [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning et benchmark réaliste.&lt;br /&gt;
* [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques pour les LLM.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix / token en temps réel.&lt;br /&gt;
* '''Qualité auto''' : précision / taux d'hallucination sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] : routage sémantique de requêtes vers les bons modèles.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache'''] : couche KV cache pour réduire la latence d'inférence.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] : orchestration de workflows critiques avec persistence.&lt;br /&gt;
* [https://github.com/lyogavin/airllm '''airLLM'''] : inférence LLM sur GPU 4 Go — pour réduire les coûts de inférence locale.&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute'''] : passerelle multi-providers (50+ gratuits), compression 15-95% de tokens — à évaluer comme fallback.&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Assistants IA &amp;amp; Agents ===&lt;br /&gt;
* [https://github.com/TencentCloud/TencentDB-Agent-Memory '''TencentDB-Agent-Memory'''] : Hub de mémoire d équipe pour agents IA - transforme conversations, docs et code en assets réutilisables.&lt;br /&gt;
* [https://github.com/unclebob/swarm-forge '''Swarm Forge'''] : Outil simple de coordination de plusieurs agents IA.&lt;br /&gt;
* [https://github.com/embabel/embabel-agent '''Embabel Agent'''] : Framework agent pour JVM (Kotlin/Java).&lt;br /&gt;
* [https://github.com/zhaoxuya520/reverse-skill '''reverse-skill'''] : Skill Router Pack pour reverse engineering et pentesting - routage IA, bootstrapping toolchain auto.&lt;br /&gt;
* [https://github.com/esengine/DeepSeek-Reasonix '''DeepSeek-Reasonix'''] : Agent de codage terminal natif DeepSeek, optimisé pour la stabilité du prefix-cache.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd'''] : Skill qui empêche l agent de coder de cacher la réponse - output clair et orienté action.&lt;br /&gt;
* [https://github.com/block/buzz '''Buzz (Block)'''] : Plateforme de communication type hive mind pour collaboration multi-agents.&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] : Transforme n importe quel PDF technique en skill Claude Code - prêt à étudier et référencer.&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/antirez/ds4 '''ds4'''] : Moteur inference local DeepSeek 4 Flash/PRO pour Metal, CUDA et ROCm - par Antirez.&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
* [https://github.com/agavra/tuicr '''tuicr''' ] - Revue de code TUI avec raccourcis vim.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2087</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2087"/>
		<updated>2026-08-07T15:15:07Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Nettoyage veilltech: suppression du catalogue GitHub, fusion avec priorités&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com/docs '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 02/06 - 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] : API de traduction multilingue temps réel — PR #3 fix timeout (gestion réponses lentes), PR #2 URLs configurables via .env.&lt;br /&gt;
* [https://github.com/ynotopec/GLM-4.7-Flash-vLLM '''GLM-4.7-Flash-vLLM'''] : Déploiement vLLM GLM-4.7 Flash.&lt;br /&gt;
* [https://github.com/ynotopec/qwen3.5-fp8-vllm '''qwen3.5-fp8-vllm'''] : Déploiement Qwen 3.5 FP8 via vLLM.&lt;br /&gt;
* [https://github.com/ynotopec/api-rag '''api-rag'''] : API RAG (recherche sémantique) — fix support FAISS SVE, suppression logs verbeux, wait-for-APIS avant build index vectoriel.&lt;br /&gt;
* [https://github.com/ynotopec/api-diarization '''api-diarization'''] : API diarisation audio — install améliorée, ergonomics auth, documentation setup projet.&lt;br /&gt;
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] : API FastAPI ASR ( Whisper ) — refonte complète, installer uv, docs runtime.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-openmed '''privacy-proxy-openmed'''] : Privacy proxy OpenMed-based — LLM upstream optionnel (mode privacy-only).&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''privacy-proxy-gliner2'''] : Privacy proxy GLiNER2-based — détection entités nommées, upstream LLM optionnel, PR #1 en cours.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] : Proxy LLM custom — paramètres modèle configurables, alias ai-default, gestion erreurs upstream.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Priorités &amp;amp; Veille ==&lt;br /&gt;
&lt;br /&gt;
Les priorités ci-dessous sont les seules entrées de veille conservées. Chaque item doit avoir un **statu** : *en cours*, *testé*, *retenu*, *abandonné*. Les projets listés sans action réelle sont retirés à chaque révision.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration d'inférence multi-nœuds.&lt;br /&gt;
* [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning et benchmark réaliste.&lt;br /&gt;
* [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques pour les LLM.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix / token en temps réel.&lt;br /&gt;
* '''Qualité auto''' : précision / taux d'hallucination sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] : routage sémantique de requêtes vers les bons modèles.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache'''] : couche KV cache pour réduire la latence d'inférence.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] : orchestration de workflows critiques avec persistence.&lt;br /&gt;
* [https://github.com/lyogavin/airllm '''airLLM'''] : inférence LLM sur GPU 4 Go — pour réduire les coûts de inférence locale.&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute'''] : passerelle multi-providers (50+ gratuits), compression 15-95% de tokens — à évaluer comme fallback.&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
* [https://github.com/agavra/tuicr '''tuicr''' ] - Revue de code TUI avec raccourcis vim.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2086</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2086"/>
		<updated>2026-08-07T01:22:55Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Weekly GitHub trending update: added embabel-agent, ds4, tuicr&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com/docs '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 02/06 - 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] : API de traduction multilingue temps réel — PR #3 fix timeout (gestion réponses lentes), PR #2 URLs configurables via .env.&lt;br /&gt;
* [https://github.com/ynotopec/GLM-4.7-Flash-vLLM '''GLM-4.7-Flash-vLLM'''] : Déploiement vLLM GLM-4.7 Flash.&lt;br /&gt;
* [https://github.com/ynotopec/qwen3.5-fp8-vllm '''qwen3.5-fp8-vllm'''] : Déploiement Qwen 3.5 FP8 via vLLM.&lt;br /&gt;
* [https://github.com/ynotopec/api-rag '''api-rag'''] : API RAG (recherche sémantique) — fix support FAISS SVE, suppression logs verbeux, wait-for-APIS avant build index vectoriel.&lt;br /&gt;
* [https://github.com/ynotopec/api-diarization '''api-diarization'''] : API diarisation audio — install améliorée, ergonomics auth, documentation setup projet.&lt;br /&gt;
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] : API FastAPI ASR ( Whisper ) — refonte complète, installer uv, docs runtime.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-openmed '''privacy-proxy-openmed'''] : Privacy proxy OpenMed-based — LLM upstream optionnel (mode privacy-only).&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''privacy-proxy-gliner2'''] : Privacy proxy GLiNER2-based — détection entités nommées, upstream LLM optionnel, PR #1 en cours.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] : Proxy LLM custom — paramètres modèle configurables, alias ai-default, gestion erreurs upstream.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/block/buzz '''buzz (Block)'''] - Plateforme de communication &amp;quot;hive mind&amp;quot; : humains et agents IA cohabitent dans les memes canaux, auto-hebergeable.&lt;br /&gt;
* [https://github.com/pingdotgg/t3code '''t3code'''] - Interface web GUI pour orchestrer les agents de codage (Codex, Claude Code, Cursor...) depuis un navigateur.&lt;br /&gt;
* [https://github.com/moeru-ai/airi '''airi (Moeru AI)'''] - Companion IA auto-heberge : voice chat temps reel, controle de jeux (Minecraft, Factorio), compatible Grok.&lt;br /&gt;
* [https://github.com/bojieli/ai-agent-book '''ai-agent-book'''] - Livre open-source (10 chapitres, 94 experiences) couvrant les AI Agents : LLM + contexte + outils, de la theorie a la production.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd''' ] - Skill pour agent de codage qui arrete le discour de cacher la reponse. Sortie ADHD-friendly.&lt;br /&gt;
* [https://github.com/agegr/pi-web '''pi-web''' ] - Interface web pour agent de codage PI avec commandes et interfaces.&lt;br /&gt;
* [https://github.com/mattpocock/skills '''skills''' ] - Skills provenant directement du repertoire .agents d'un ingenieur reelement experimente.&lt;br /&gt;
* [https://github.com/earthtojake/text-to-cad '''text-to-cad''' ] - Collection de skills agent pour CAO, robotique et conception hardware.&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] - Transforme tout PDF technique en competence Claude Code : pret a etudier, referencer et utiliser pendant le travail.. (13726 stars)&lt;br /&gt;
* [https://github.com/UditAkhourii/adhd '''adhd'''] - Skill agent de codage : tree-of-thought avec pruning, base sur Claude &amp;amp; Codex Agent SDK. Ideation divergente.. (2879 stars)&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/different-ai/openwork openwork] - Alternative open-source a Claude Cowork (base sur opencode) : harness d agents de codage auto-heberge, multi-agents et multi-modeles.* [https://github.com/citrolabs/ego-lite '''ego-lite'''] - Navigateur optimise pour les agents IA : automation avec etat de session, compatible Codex et Claude Code.. (6536 stars)&lt;br /&gt;
* [https://github.com/zhaoxuya520/reverse-skill '''reverse-skill''' ] - Pack de compétences IA avec routage intelligent pour reverse engineering, pentesting autorisé et recherche de sécurité.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/TencentCloud/TencentDB-Agent-Memory '''TencentDB-Agent-Memory''' ] — Hub de mémoire d'équipe pour agents IA : transforme conversations, docs et code en 4 actifs réutilisables (Chat Memory, Skill, LLM-Wiki, Code-Graph).&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/esengine/DeepSeek-Reasonix '''DeepSeek-Reasonix''' ] - Agent de codage IA natif DeepSeek : optimise le prefix-cache pour une execution stable en terminal, peut tourner en continu.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/citrolabs/ego-lite '''ego-lite''' ] - Navigateur ultra-rapide pour agents IA : partagez votre session navigateurs via MCP, compatible Codex et Claude Code. 0 cout, 0 config.&lt;br /&gt;
* [https://github.com/embabel/embabel-agent '''embabel-agent''' ] - Framework d'agents IA pour JVM. Prononce Em-BAY-bel /mbeibel/.&lt;br /&gt;
&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Superto''']&lt;br /&gt;
* [https://github.com/microsoft/VibeVoice '''VibeVoice (Microsoft)'''] - Framework vocal open-source (TTS + ASR) de Microsoft : synthese, clonage et reconnaissance vocale de '''pointe.nic-3''' — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/andrewyng/aisuite '''aisuite'''] - Interface unifiee pour les LLM : un seul appel OpenAI-compatible pour OpenAI, Anthropic, Google, AWS, Cohere, Ollama, OpenRouter...&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les L&lt;br /&gt;
* [https://github.com/opengeos/GeoLibre GeoLibre] - Plateforme SIG open-source cloud-native pour visualisation et analyse de donnees geospatiales (web, bureau, mobile, Jupyter).&lt;br /&gt;
* [https://github.com/CoreBunch/Instatic Instatic] - CMS visuel auto-heberge et agentic alternative a Webflow/Framer : utilisateurs, roles, plugins, base de donnees integree.LM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
* [https://github.com/lyogavin/airllm '''airllm''' ] - Inférence LLM haute performance sur un seul GPU 4Go, permettant l'exécution de modèles de 70B+ paramètres en local.&lt;br /&gt;
* [https://github.com/oblien/openship '''openship''' ] - Plateforme de deploiement auto-hebergeee et auto-deployee avec travaux automatiques.&lt;br /&gt;
* [https://github.com/antirez/ds4 '''ds4''' ] - Moteur d'inférence local DeepSeek 4 Flash et PRO pour Metal, CUDA et ROCm.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://github.com/KnockOutEZ/wigolo '''wigolo''' ] - Web local-first pour agent de codage — recherche, fetch, crawl et recherche via MCP, sans clee API.&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
* [https://github.com/agavra/tuicr '''tuicr''' ] - Revue de code TUI avec raccourcis vim.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2085</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2085"/>
		<updated>2026-08-06T12:44:45Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Undo revision 2082 by Tcepo (talk)&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com/docs '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 02/06 - 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] : API de traduction multilingue temps réel — PR #3 fix timeout (gestion réponses lentes), PR #2 URLs configurables via .env.&lt;br /&gt;
* [https://github.com/ynotopec/GLM-4.7-Flash-vLLM '''GLM-4.7-Flash-vLLM'''] : Déploiement vLLM GLM-4.7 Flash.&lt;br /&gt;
* [https://github.com/ynotopec/qwen3.5-fp8-vllm '''qwen3.5-fp8-vllm'''] : Déploiement Qwen 3.5 FP8 via vLLM.&lt;br /&gt;
* [https://github.com/ynotopec/api-rag '''api-rag'''] : API RAG (recherche sémantique) — fix support FAISS SVE, suppression logs verbeux, wait-for-APIS avant build index vectoriel.&lt;br /&gt;
* [https://github.com/ynotopec/api-diarization '''api-diarization'''] : API diarisation audio — install améliorée, ergonomics auth, documentation setup projet.&lt;br /&gt;
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] : API FastAPI ASR ( Whisper ) — refonte complète, installer uv, docs runtime.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-openmed '''privacy-proxy-openmed'''] : Privacy proxy OpenMed-based — LLM upstream optionnel (mode privacy-only).&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''privacy-proxy-gliner2'''] : Privacy proxy GLiNER2-based — détection entités nommées, upstream LLM optionnel, PR #1 en cours.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] : Proxy LLM custom — paramètres modèle configurables, alias ai-default, gestion erreurs upstream.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/block/buzz '''buzz (Block)'''] - Plateforme de communication &amp;quot;hive mind&amp;quot; : humains et agents IA cohabitent dans les memes canaux, auto-hebergeable.&lt;br /&gt;
* [https://github.com/pingdotgg/t3code '''t3code'''] - Interface web GUI pour orchestrer les agents de codage (Codex, Claude Code, Cursor...) depuis un navigateur.&lt;br /&gt;
* [https://github.com/moeru-ai/airi '''airi (Moeru AI)'''] - Companion IA auto-heberge : voice chat temps reel, controle de jeux (Minecraft, Factorio), compatible Grok.&lt;br /&gt;
* [https://github.com/bojieli/ai-agent-book '''ai-agent-book'''] - Livre open-source (10 chapitres, 94 experiences) couvrant les AI Agents : LLM + contexte + outils, de la theorie a la production.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd''' ] - Skill pour agent de codage qui arrete le discour de cacher la reponse. Sortie ADHD-friendly.&lt;br /&gt;
* [https://github.com/agegr/pi-web '''pi-web''' ] - Interface web pour agent de codage PI avec commandes et interfaces.&lt;br /&gt;
* [https://github.com/mattpocock/skills '''skills''' ] - Skills provenant directement du repertoire .agents d'un ingenieur reelement experimente.&lt;br /&gt;
* [https://github.com/earthtojake/text-to-cad '''text-to-cad''' ] - Collection de skills agent pour CAO, robotique et conception hardware.&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] - Transforme tout PDF technique en competence Claude Code : pret a etudier, referencer et utiliser pendant le travail.. (13726 stars)&lt;br /&gt;
* [https://github.com/UditAkhourii/adhd '''adhd'''] - Skill agent de codage : tree-of-thought avec pruning, base sur Claude &amp;amp; Codex Agent SDK. Ideation divergente.. (2879 stars)&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/different-ai/openwork openwork] - Alternative open-source a Claude Cowork (base sur opencode) : harness d agents de codage auto-heberge, multi-agents et multi-modeles.* [https://github.com/citrolabs/ego-lite '''ego-lite'''] - Navigateur optimise pour les agents IA : automation avec etat de session, compatible Codex et Claude Code.. (6536 stars)&lt;br /&gt;
* [https://github.com/zhaoxuya520/reverse-skill '''reverse-skill''' ] - Pack de compétences IA avec routage intelligent pour reverse engineering, pentesting autorisé et recherche de sécurité.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/TencentCloud/TencentDB-Agent-Memory '''TencentDB-Agent-Memory''' ] — Hub de mémoire d'équipe pour agents IA : transforme conversations, docs et code en 4 actifs réutilisables (Chat Memory, Skill, LLM-Wiki, Code-Graph).&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/esengine/DeepSeek-Reasonix '''DeepSeek-Reasonix''' ] - Agent de codage IA natif DeepSeek : optimise le prefix-cache pour une execution stable en terminal, peut tourner en continu.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/citrolabs/ego-lite '''ego-lite''' ] - Navigateur ultra-rapide pour agents IA : partagez votre session navigateurs via MCP, compatible Codex et Claude Code. 0 cout, 0 config.&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Superto''']&lt;br /&gt;
* [https://github.com/microsoft/VibeVoice '''VibeVoice (Microsoft)'''] - Framework vocal open-source (TTS + ASR) de Microsoft : synthese, clonage et reconnaissance vocale de '''pointe.nic-3''' — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/andrewyng/aisuite '''aisuite'''] - Interface unifiee pour les LLM : un seul appel OpenAI-compatible pour OpenAI, Anthropic, Google, AWS, Cohere, Ollama, OpenRouter...&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les L&lt;br /&gt;
* [https://github.com/opengeos/GeoLibre GeoLibre] - Plateforme SIG open-source cloud-native pour visualisation et analyse de donnees geospatiales (web, bureau, mobile, Jupyter).&lt;br /&gt;
* [https://github.com/CoreBunch/Instatic Instatic] - CMS visuel auto-heberge et agentic alternative a Webflow/Framer : utilisateurs, roles, plugins, base de donnees integree.LM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
* [https://github.com/lyogavin/airllm '''airllm''' ] - Inférence LLM haute performance sur un seul GPU 4Go, permettant l'exécution de modèles de 70B+ paramètres en local.&lt;br /&gt;
* [https://github.com/oblien/openship '''openship''' ] - Plateforme de deploiement auto-hebergeee et auto-deployee avec travaux automatiques.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://github.com/KnockOutEZ/wigolo '''wigolo''' ] - Web local-first pour agent de codage — recherche, fetch, crawl et recherche via MCP, sans clee API.&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2084</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2084"/>
		<updated>2026-08-06T12:44:27Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Undo revision 2083 by Tcepo (talk)&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] : API de traduction multilingue temps réel — PR #3 fix timeout (gestion réponses lentes), PR #2 URLs configurables via .env.&lt;br /&gt;
* [https://github.com/ynotopec/GLM-4.7-Flash-vLLM '''GLM-4.7-Flash-vLLM'''] : Déploiement vLLM GLM-4.7 Flash.&lt;br /&gt;
* [https://github.com/ynotopec/qwen3.5-fp8-vllm '''qwen3.5-fp8-vllm'''] : Déploiement Qwen 3.5 FP8 via vLLM.&lt;br /&gt;
* [https://github.com/ynotopec/api-rag '''api-rag'''] : API RAG (recherche sémantique) — fix support FAISS SVE, suppression logs verbeux, wait-for-APIS avant build index vectoriel.&lt;br /&gt;
* [https://github.com/ynotopec/api-diarization '''api-diarization'''] : API diarisation audio — install améliorée, ergonomics auth, documentation setup projet.&lt;br /&gt;
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] : API FastAPI ASR ( Whisper ) — refonte complète, installer uv, docs runtime.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-openmed '''privacy-proxy-openmed'''] : Privacy proxy OpenMed-based — LLM upstream optionnel (mode privacy-only).&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''privacy-proxy-gliner2'''] : Privacy proxy GLiNER2-based — détection entités nommées, upstream LLM optionnel, PR #1 en cours.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] : Proxy LLM custom — paramètres modèle configurables, alias ai-default, gestion erreurs upstream.&lt;br /&gt;
&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com/docs '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/block/buzz '''buzz (Block)'''] - Plateforme de communication &amp;quot;hive mind&amp;quot; : humains et agents IA cohabitent dans les memes canaux, auto-hebergeable.&lt;br /&gt;
* [https://github.com/pingdotgg/t3code '''t3code'''] - Interface web GUI pour orchestrer les agents de codage (Codex, Claude Code, Cursor...) depuis un navigateur.&lt;br /&gt;
* [https://github.com/moeru-ai/airi '''airi (Moeru AI)'''] - Companion IA auto-heberge : voice chat temps reel, controle de jeux (Minecraft, Factorio), compatible Grok.&lt;br /&gt;
* [https://github.com/bojieli/ai-agent-book '''ai-agent-book'''] - Livre open-source (10 chapitres, 94 experiences) couvrant les AI Agents : LLM + contexte + outils, de la theorie a la production.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd''' ] - Skill pour agent de codage qui arrete le discour de cacher la reponse. Sortie ADHD-friendly.&lt;br /&gt;
* [https://github.com/agegr/pi-web '''pi-web''' ] - Interface web pour agent de codage PI avec commandes et interfaces.&lt;br /&gt;
* [https://github.com/mattpocock/skills '''skills''' ] - Skills provenant directement du repertoire .agents d'un ingenieur reelement experimente.&lt;br /&gt;
* [https://github.com/earthtojake/text-to-cad '''text-to-cad''' ] - Collection de skills agent pour CAO, robotique et conception hardware.&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] - Transforme tout PDF technique en competence Claude Code : pret a etudier, referencer et utiliser pendant le travail.. (13726 stars)&lt;br /&gt;
* [https://github.com/UditAkhourii/adhd '''adhd'''] - Skill agent de codage : tree-of-thought avec pruning, base sur Claude &amp;amp; Codex Agent SDK. Ideation divergente.. (2879 stars)&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/different-ai/openwork openwork] - Alternative open-source a Claude Cowork (base sur opencode) : harness d agents de codage auto-heberge, multi-agents et multi-modeles.* [https://github.com/citrolabs/ego-lite '''ego-lite'''] - Navigateur optimise pour les agents IA : automation avec etat de session, compatible Codex et Claude Code.. (6536 stars)&lt;br /&gt;
* [https://github.com/zhaoxuya520/reverse-skill '''reverse-skill''' ] - Pack de compétences IA avec routage intelligent pour reverse engineering, pentesting autorisé et recherche de sécurité.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/TencentCloud/TencentDB-Agent-Memory '''TencentDB-Agent-Memory''' ] — Hub de mémoire d'équipe pour agents IA : transforme conversations, docs et code en 4 actifs réutilisables (Chat Memory, Skill, LLM-Wiki, Code-Graph).&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/esengine/DeepSeek-Reasonix '''DeepSeek-Reasonix''' ] - Agent de codage IA natif DeepSeek : optimise le prefix-cache pour une execution stable en terminal, peut tourner en continu.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/citrolabs/ego-lite '''ego-lite''' ] - Navigateur ultra-rapide pour agents IA : partagez votre session navigateurs via MCP, compatible Codex et Claude Code. 0 cout, 0 config.&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Superto''']&lt;br /&gt;
* [https://github.com/microsoft/VibeVoice '''VibeVoice (Microsoft)'''] - Framework vocal open-source (TTS + ASR) de Microsoft : synthese, clonage et reconnaissance vocale de '''pointe.nic-3''' — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/andrewyng/aisuite '''aisuite'''] - Interface unifiee pour les LLM : un seul appel OpenAI-compatible pour OpenAI, Anthropic, Google, AWS, Cohere, Ollama, OpenRouter...&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les L&lt;br /&gt;
* [https://github.com/opengeos/GeoLibre GeoLibre] - Plateforme SIG open-source cloud-native pour visualisation et analyse de donnees geospatiales (web, bureau, mobile, Jupyter).&lt;br /&gt;
* [https://github.com/CoreBunch/Instatic Instatic] - CMS visuel auto-heberge et agentic alternative a Webflow/Framer : utilisateurs, roles, plugins, base de donnees integree.LM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
* [https://github.com/lyogavin/airllm '''airllm''' ] - Inférence LLM haute performance sur un seul GPU 4Go, permettant l'exécution de modèles de 70B+ paramètres en local.&lt;br /&gt;
* [https://github.com/oblien/openship '''openship''' ] - Plateforme de deploiement auto-hebergeee et auto-deployee avec travaux automatiques.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://github.com/KnockOutEZ/wigolo '''wigolo''' ] - Web local-first pour agent de codage — recherche, fetch, crawl et recherche via MCP, sans clee API.&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2083</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2083"/>
		<updated>2026-08-06T12:42:21Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Reverted edits by Tcepo (talk) to last revision by Kosstradukso&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Admin UNIX==&lt;br /&gt;
*[http://www.unixguide.net/unixguide.shtml UNIX Guide]&lt;br /&gt;
*[http://wiki.gantzer.eu/index.php/Notes_en_vrac_(brouillon) Patrol settings]&lt;br /&gt;
I actually knew about most of this, but having said that, I still thought it was useful. Nice job!&lt;br /&gt;
Jean de La Fontaine~ A cheerful mind is a vigorous mind. [http://www.medicinaa.com buy viagra online uk no prescription]&lt;br /&gt;
&lt;br /&gt;
== Today keys ==&lt;br /&gt;
&lt;br /&gt;
*SNOW 320 (IP phone)&lt;br /&gt;
&lt;br /&gt;
*SIPXECS (Call server, free)&lt;br /&gt;
&lt;br /&gt;
*MEDIAWIKI (Shared editor, free)&lt;br /&gt;
&lt;br /&gt;
*UBUNTU SERVER (OS, free)&lt;br /&gt;
&lt;br /&gt;
*PHPBB (Forum, free)&lt;br /&gt;
&lt;br /&gt;
*BUGZILLA (Bug management, free)&lt;br /&gt;
&lt;br /&gt;
*ALFRESCO (Information management, free)&lt;br /&gt;
&lt;br /&gt;
*SME (Company management, free)&lt;br /&gt;
&lt;br /&gt;
*SOLARIS (Data center)&lt;br /&gt;
&lt;br /&gt;
*BUSYBOX (Embedded system, free)&lt;br /&gt;
&lt;br /&gt;
*VMWARE (Infrastucture)&lt;br /&gt;
&lt;br /&gt;
*OTRS (Desk management, free)&lt;br /&gt;
&lt;br /&gt;
== Administration tools ==&lt;br /&gt;
&lt;br /&gt;
* partimage (Data backup)&lt;br /&gt;
&lt;br /&gt;
* tcpdump (Network monitoring)&lt;br /&gt;
&lt;br /&gt;
* nmap (Network scanner)&lt;br /&gt;
&lt;br /&gt;
* airodum-ng (Wireless monitoring)&lt;br /&gt;
&lt;br /&gt;
* IBM/Hitachi's Drive Fitness Test&lt;br /&gt;
&lt;br /&gt;
* SiSoft Sandra, Mainboard Monitor ; CPU test&lt;br /&gt;
&lt;br /&gt;
== Hosting ==&lt;br /&gt;
&lt;br /&gt;
* [http://www.eapps.com eapps.com]&lt;br /&gt;
* [http://www.hesk.com hesk.com]&lt;br /&gt;
&lt;br /&gt;
== Technical model ==&lt;br /&gt;
&lt;br /&gt;
* [http://wikitech.wikimedia.org/view/Main_Page Wikipedia]&lt;br /&gt;
&lt;br /&gt;
== Training ==&lt;br /&gt;
&lt;br /&gt;
*[http://www.gocertify.com/quizzes/ccna/ccna2_19.shtml NETWORK (CCNA)]&lt;br /&gt;
&lt;br /&gt;
*[http://www.vmuser.com/joomla/index.php?option=com_joomlaquiz&amp;amp;Itemid=7 VMWARE]&lt;br /&gt;
&lt;br /&gt;
*ITIL v2 - Foundations&lt;br /&gt;
&lt;br /&gt;
== Links ==&lt;br /&gt;
&lt;br /&gt;
*[http://jobsearchtech.about.com/od/educationfortechcareers/tp/HighestCerts.htm Top technical certifications]&lt;br /&gt;
&lt;br /&gt;
==Salaries==&lt;br /&gt;
&lt;br /&gt;
*[http://www.insee.fr/fr/themes/tableau.asp?reg_id=0&amp;amp;ref_id=NATSEF04119 INSEE 2008]&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2082</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2082"/>
		<updated>2026-08-06T12:41:20Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] : API de traduction multilingue temps réel — PR #3 fix timeout (gestion réponses lentes), PR #2 URLs configurables via .env.&lt;br /&gt;
* [https://github.com/ynotopec/GLM-4.7-Flash-vLLM '''GLM-4.7-Flash-vLLM'''] : Déploiement vLLM GLM-4.7 Flash.&lt;br /&gt;
* [https://github.com/ynotopec/qwen3.5-fp8-vllm '''qwen3.5-fp8-vllm'''] : Déploiement Qwen 3.5 FP8 via vLLM.&lt;br /&gt;
* [https://github.com/ynotopec/api-rag '''api-rag'''] : API RAG (recherche sémantique) — fix support FAISS SVE, suppression logs verbeux, wait-for-APIS avant build index vectoriel.&lt;br /&gt;
* [https://github.com/ynotopec/api-diarization '''api-diarization'''] : API diarisation audio — install améliorée, ergonomics auth, documentation setup projet.&lt;br /&gt;
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] : API FastAPI ASR ( Whisper ) — refonte complète, installer uv, docs runtime.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-openmed '''privacy-proxy-openmed'''] : Privacy proxy OpenMed-based — LLM upstream optionnel (mode privacy-only).&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''privacy-proxy-gliner2'''] : Privacy proxy GLiNER2-based — détection entités nommées, upstream LLM optionnel, PR #1 en cours.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] : Proxy LLM custom — paramètres modèle configurables, alias ai-default, gestion erreurs upstream.&lt;br /&gt;
&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com/docs '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/block/buzz '''buzz (Block)'''] - Plateforme de communication &amp;quot;hive mind&amp;quot; : humains et agents IA cohabitent dans les memes canaux, auto-hebergeable.&lt;br /&gt;
* [https://github.com/pingdotgg/t3code '''t3code'''] - Interface web GUI pour orchestrer les agents de codage (Codex, Claude Code, Cursor...) depuis un navigateur.&lt;br /&gt;
* [https://github.com/moeru-ai/airi '''airi (Moeru AI)'''] - Companion IA auto-heberge : voice chat temps reel, controle de jeux (Minecraft, Factorio), compatible Grok.&lt;br /&gt;
* [https://github.com/bojieli/ai-agent-book '''ai-agent-book'''] - Livre open-source (10 chapitres, 94 experiences) couvrant les AI Agents : LLM + contexte + outils, de la theorie a la production.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd''' ] - Skill pour agent de codage qui arrete le discour de cacher la reponse. Sortie ADHD-friendly.&lt;br /&gt;
* [https://github.com/agegr/pi-web '''pi-web''' ] - Interface web pour agent de codage PI avec commandes et interfaces.&lt;br /&gt;
* [https://github.com/mattpocock/skills '''skills''' ] - Skills provenant directement du repertoire .agents d'un ingenieur reelement experimente.&lt;br /&gt;
* [https://github.com/earthtojake/text-to-cad '''text-to-cad''' ] - Collection de skills agent pour CAO, robotique et conception hardware.&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] - Transforme tout PDF technique en competence Claude Code : pret a etudier, referencer et utiliser pendant le travail.. (13726 stars)&lt;br /&gt;
* [https://github.com/UditAkhourii/adhd '''adhd'''] - Skill agent de codage : tree-of-thought avec pruning, base sur Claude &amp;amp; Codex Agent SDK. Ideation divergente.. (2879 stars)&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/different-ai/openwork openwork] - Alternative open-source a Claude Cowork (base sur opencode) : harness d agents de codage auto-heberge, multi-agents et multi-modeles.* [https://github.com/citrolabs/ego-lite '''ego-lite'''] - Navigateur optimise pour les agents IA : automation avec etat de session, compatible Codex et Claude Code.. (6536 stars)&lt;br /&gt;
* [https://github.com/zhaoxuya520/reverse-skill '''reverse-skill''' ] - Pack de compétences IA avec routage intelligent pour reverse engineering, pentesting autorisé et recherche de sécurité.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/TencentCloud/TencentDB-Agent-Memory '''TencentDB-Agent-Memory''' ] — Hub de mémoire d'équipe pour agents IA : transforme conversations, docs et code en 4 actifs réutilisables (Chat Memory, Skill, LLM-Wiki, Code-Graph).&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/esengine/DeepSeek-Reasonix '''DeepSeek-Reasonix''' ] - Agent de codage IA natif DeepSeek : optimise le prefix-cache pour une execution stable en terminal, peut tourner en continu.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/citrolabs/ego-lite '''ego-lite''' ] - Navigateur ultra-rapide pour agents IA : partagez votre session navigateurs via MCP, compatible Codex et Claude Code. 0 cout, 0 config.&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Superto''']&lt;br /&gt;
* [https://github.com/microsoft/VibeVoice '''VibeVoice (Microsoft)'''] - Framework vocal open-source (TTS + ASR) de Microsoft : synthese, clonage et reconnaissance vocale de '''pointe.nic-3''' — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/andrewyng/aisuite '''aisuite'''] - Interface unifiee pour les LLM : un seul appel OpenAI-compatible pour OpenAI, Anthropic, Google, AWS, Cohere, Ollama, OpenRouter...&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les L&lt;br /&gt;
* [https://github.com/opengeos/GeoLibre GeoLibre] - Plateforme SIG open-source cloud-native pour visualisation et analyse de donnees geospatiales (web, bureau, mobile, Jupyter).&lt;br /&gt;
* [https://github.com/CoreBunch/Instatic Instatic] - CMS visuel auto-heberge et agentic alternative a Webflow/Framer : utilisateurs, roles, plugins, base de donnees integree.LM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
* [https://github.com/lyogavin/airllm '''airllm''' ] - Inférence LLM haute performance sur un seul GPU 4Go, permettant l'exécution de modèles de 70B+ paramètres en local.&lt;br /&gt;
* [https://github.com/oblien/openship '''openship''' ] - Plateforme de deploiement auto-hebergeee et auto-deployee avec travaux automatiques.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://github.com/KnockOutEZ/wigolo '''wigolo''' ] - Web local-first pour agent de codage — recherche, fetch, crawl et recherche via MCP, sans clee API.&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2081</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2081"/>
		<updated>2026-08-06T09:14:37Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: /* Exemple */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com/docs '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 02/06 - 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] : API de traduction multilingue temps réel — PR #3 fix timeout (gestion réponses lentes), PR #2 URLs configurables via .env.&lt;br /&gt;
* [https://github.com/ynotopec/GLM-4.7-Flash-vLLM '''GLM-4.7-Flash-vLLM'''] : Déploiement vLLM GLM-4.7 Flash.&lt;br /&gt;
* [https://github.com/ynotopec/qwen3.5-fp8-vllm '''qwen3.5-fp8-vllm'''] : Déploiement Qwen 3.5 FP8 via vLLM.&lt;br /&gt;
* [https://github.com/ynotopec/api-rag '''api-rag'''] : API RAG (recherche sémantique) — fix support FAISS SVE, suppression logs verbeux, wait-for-APIS avant build index vectoriel.&lt;br /&gt;
* [https://github.com/ynotopec/api-diarization '''api-diarization'''] : API diarisation audio — install améliorée, ergonomics auth, documentation setup projet.&lt;br /&gt;
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] : API FastAPI ASR ( Whisper ) — refonte complète, installer uv, docs runtime.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-openmed '''privacy-proxy-openmed'''] : Privacy proxy OpenMed-based — LLM upstream optionnel (mode privacy-only).&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''privacy-proxy-gliner2'''] : Privacy proxy GLiNER2-based — détection entités nommées, upstream LLM optionnel, PR #1 en cours.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] : Proxy LLM custom — paramètres modèle configurables, alias ai-default, gestion erreurs upstream.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/block/buzz '''buzz (Block)'''] - Plateforme de communication &amp;quot;hive mind&amp;quot; : humains et agents IA cohabitent dans les memes canaux, auto-hebergeable.&lt;br /&gt;
* [https://github.com/pingdotgg/t3code '''t3code'''] - Interface web GUI pour orchestrer les agents de codage (Codex, Claude Code, Cursor...) depuis un navigateur.&lt;br /&gt;
* [https://github.com/moeru-ai/airi '''airi (Moeru AI)'''] - Companion IA auto-heberge : voice chat temps reel, controle de jeux (Minecraft, Factorio), compatible Grok.&lt;br /&gt;
* [https://github.com/bojieli/ai-agent-book '''ai-agent-book'''] - Livre open-source (10 chapitres, 94 experiences) couvrant les AI Agents : LLM + contexte + outils, de la theorie a la production.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd''' ] - Skill pour agent de codage qui arrete le discour de cacher la reponse. Sortie ADHD-friendly.&lt;br /&gt;
* [https://github.com/agegr/pi-web '''pi-web''' ] - Interface web pour agent de codage PI avec commandes et interfaces.&lt;br /&gt;
* [https://github.com/mattpocock/skills '''skills''' ] - Skills provenant directement du repertoire .agents d'un ingenieur reelement experimente.&lt;br /&gt;
* [https://github.com/earthtojake/text-to-cad '''text-to-cad''' ] - Collection de skills agent pour CAO, robotique et conception hardware.&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] - Transforme tout PDF technique en competence Claude Code : pret a etudier, referencer et utiliser pendant le travail.. (13726 stars)&lt;br /&gt;
* [https://github.com/UditAkhourii/adhd '''adhd'''] - Skill agent de codage : tree-of-thought avec pruning, base sur Claude &amp;amp; Codex Agent SDK. Ideation divergente.. (2879 stars)&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/different-ai/openwork openwork] - Alternative open-source a Claude Cowork (base sur opencode) : harness d agents de codage auto-heberge, multi-agents et multi-modeles.* [https://github.com/citrolabs/ego-lite '''ego-lite'''] - Navigateur optimise pour les agents IA : automation avec etat de session, compatible Codex et Claude Code.. (6536 stars)&lt;br /&gt;
* [https://github.com/zhaoxuya520/reverse-skill '''reverse-skill''' ] - Pack de compétences IA avec routage intelligent pour reverse engineering, pentesting autorisé et recherche de sécurité.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/TencentCloud/TencentDB-Agent-Memory '''TencentDB-Agent-Memory''' ] — Hub de mémoire d'équipe pour agents IA : transforme conversations, docs et code en 4 actifs réutilisables (Chat Memory, Skill, LLM-Wiki, Code-Graph).&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/esengine/DeepSeek-Reasonix '''DeepSeek-Reasonix''' ] - Agent de codage IA natif DeepSeek : optimise le prefix-cache pour une execution stable en terminal, peut tourner en continu.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/citrolabs/ego-lite '''ego-lite''' ] - Navigateur ultra-rapide pour agents IA : partagez votre session navigateurs via MCP, compatible Codex et Claude Code. 0 cout, 0 config.&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Superto''']&lt;br /&gt;
* [https://github.com/microsoft/VibeVoice '''VibeVoice (Microsoft)'''] - Framework vocal open-source (TTS + ASR) de Microsoft : synthese, clonage et reconnaissance vocale de '''pointe.nic-3''' — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/andrewyng/aisuite '''aisuite'''] - Interface unifiee pour les LLM : un seul appel OpenAI-compatible pour OpenAI, Anthropic, Google, AWS, Cohere, Ollama, OpenRouter...&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les L&lt;br /&gt;
* [https://github.com/opengeos/GeoLibre GeoLibre] - Plateforme SIG open-source cloud-native pour visualisation et analyse de donnees geospatiales (web, bureau, mobile, Jupyter).&lt;br /&gt;
* [https://github.com/CoreBunch/Instatic Instatic] - CMS visuel auto-heberge et agentic alternative a Webflow/Framer : utilisateurs, roles, plugins, base de donnees integree.LM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
* [https://github.com/lyogavin/airllm '''airllm''' ] - Inférence LLM haute performance sur un seul GPU 4Go, permettant l'exécution de modèles de 70B+ paramètres en local.&lt;br /&gt;
* [https://github.com/oblien/openship '''openship''' ] - Plateforme de deploiement auto-hebergeee et auto-deployee avec travaux automatiques.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://github.com/KnockOutEZ/wigolo '''wigolo''' ] - Web local-first pour agent de codage — recherche, fetch, crawl et recherche via MCP, sans clee API.&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2080</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2080"/>
		<updated>2026-08-06T01:21:08Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Ajout DeepSeek-Reasonix et ego-lite - Weekly Trending&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com/docs '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 02/06 - 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] : API de traduction multilingue temps réel — PR #3 fix timeout (gestion réponses lentes), PR #2 URLs configurables via .env.&lt;br /&gt;
* [https://github.com/ynotopec/GLM-4.7-Flash-vLLM '''GLM-4.7-Flash-vLLM'''] : Déploiement vLLM GLM-4.7 Flash.&lt;br /&gt;
* [https://github.com/ynotopec/qwen3.5-fp8-vllm '''qwen3.5-fp8-vllm'''] : Déploiement Qwen 3.5 FP8 via vLLM.&lt;br /&gt;
* [https://github.com/ynotopec/api-rag '''api-rag'''] : API RAG (recherche sémantique) — fix support FAISS SVE, suppression logs verbeux, wait-for-APIS avant build index vectoriel.&lt;br /&gt;
* [https://github.com/ynotopec/api-diarization '''api-diarization'''] : API diarisation audio — install améliorée, ergonomics auth, documentation setup projet.&lt;br /&gt;
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] : API FastAPI ASR ( Whisper ) — refonte complète, installer uv, docs runtime.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-openmed '''privacy-proxy-openmed'''] : Privacy proxy OpenMed-based — LLM upstream optionnel (mode privacy-only).&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''privacy-proxy-gliner2'''] : Privacy proxy GLiNER2-based — détection entités nommées, upstream LLM optionnel, PR #1 en cours.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] : Proxy LLM custom — paramètres modèle configurables, alias ai-default, gestion erreurs upstream.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/block/buzz '''buzz (Block)'''] - Plateforme de communication &amp;quot;hive mind&amp;quot; : humains et agents IA cohabitent dans les memes canaux, auto-hebergeable.&lt;br /&gt;
* [https://github.com/pingdotgg/t3code '''t3code'''] - Interface web GUI pour orchestrer les agents de codage (Codex, Claude Code, Cursor...) depuis un navigateur.&lt;br /&gt;
* [https://github.com/moeru-ai/airi '''airi (Moeru AI)'''] - Companion IA auto-heberge : voice chat temps reel, controle de jeux (Minecraft, Factorio), compatible Grok.&lt;br /&gt;
* [https://github.com/bojieli/ai-agent-book '''ai-agent-book'''] - Livre open-source (10 chapitres, 94 experiences) couvrant les AI Agents : LLM + contexte + outils, de la theorie a la production.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd''' ] - Skill pour agent de codage qui arrete le discour de cacher la reponse. Sortie ADHD-friendly.&lt;br /&gt;
* [https://github.com/agegr/pi-web '''pi-web''' ] - Interface web pour agent de codage PI avec commandes et interfaces.&lt;br /&gt;
* [https://github.com/mattpocock/skills '''skills''' ] - Skills provenant directement du repertoire .agents d'un ingenieur reelement experimente.&lt;br /&gt;
* [https://github.com/earthtojake/text-to-cad '''text-to-cad''' ] - Collection de skills agent pour CAO, robotique et conception hardware.&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] - Transforme tout PDF technique en competence Claude Code : pret a etudier, referencer et utiliser pendant le travail.. (13726 stars)&lt;br /&gt;
* [https://github.com/UditAkhourii/adhd '''adhd'''] - Skill agent de codage : tree-of-thought avec pruning, base sur Claude &amp;amp; Codex Agent SDK. Ideation divergente.. (2879 stars)&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/different-ai/openwork openwork] - Alternative open-source a Claude Cowork (base sur opencode) : harness d agents de codage auto-heberge, multi-agents et multi-modeles.* [https://github.com/citrolabs/ego-lite '''ego-lite'''] - Navigateur optimise pour les agents IA : automation avec etat de session, compatible Codex et Claude Code.. (6536 stars)&lt;br /&gt;
* [https://github.com/zhaoxuya520/reverse-skill '''reverse-skill''' ] - Pack de compétences IA avec routage intelligent pour reverse engineering, pentesting autorisé et recherche de sécurité.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/TencentCloud/TencentDB-Agent-Memory '''TencentDB-Agent-Memory''' ] — Hub de mémoire d'équipe pour agents IA : transforme conversations, docs et code en 4 actifs réutilisables (Chat Memory, Skill, LLM-Wiki, Code-Graph).&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/esengine/DeepSeek-Reasonix '''DeepSeek-Reasonix''' ] - Agent de codage IA natif DeepSeek : optimise le prefix-cache pour une execution stable en terminal, peut tourner en continu.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/citrolabs/ego-lite '''ego-lite''' ] - Navigateur ultra-rapide pour agents IA : partagez votre session navigateurs via MCP, compatible Codex et Claude Code. 0 cout, 0 config.&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Superto''']&lt;br /&gt;
* [https://github.com/microsoft/VibeVoice '''VibeVoice (Microsoft)'''] - Framework vocal open-source (TTS + ASR) de Microsoft : synthese, clonage et reconnaissance vocale de '''pointe.nic-3''' — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/andrewyng/aisuite '''aisuite'''] - Interface unifiee pour les LLM : un seul appel OpenAI-compatible pour OpenAI, Anthropic, Google, AWS, Cohere, Ollama, OpenRouter...&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les L&lt;br /&gt;
* [https://github.com/opengeos/GeoLibre GeoLibre] - Plateforme SIG open-source cloud-native pour visualisation et analyse de donnees geospatiales (web, bureau, mobile, Jupyter).&lt;br /&gt;
* [https://github.com/CoreBunch/Instatic Instatic] - CMS visuel auto-heberge et agentic alternative a Webflow/Framer : utilisateurs, roles, plugins, base de donnees integree.LM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
* [https://github.com/lyogavin/airllm '''airllm''' ] - Inférence LLM haute performance sur un seul GPU 4Go, permettant l'exécution de modèles de 70B+ paramètres en local.&lt;br /&gt;
* [https://github.com/oblien/openship '''openship''' ] - Plateforme de deploiement auto-hebergeee et auto-deployee avec travaux automatiques.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://github.com/KnockOutEZ/wigolo '''wigolo''' ] - Web local-first pour agent de codage — recherche, fetch, crawl et recherche via MCP, sans clee API.&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2079</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2079"/>
		<updated>2026-08-06T01:17:58Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com/docs '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 02/06 - 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] : API de traduction multilingue temps réel — PR #3 fix timeout (gestion réponses lentes), PR #2 URLs configurables via .env.&lt;br /&gt;
* [https://github.com/ynotopec/GLM-4.7-Flash-vLLM '''GLM-4.7-Flash-vLLM'''] : Déploiement vLLM GLM-4.7 Flash.&lt;br /&gt;
* [https://github.com/ynotopec/qwen3.5-fp8-vllm '''qwen3.5-fp8-vllm'''] : Déploiement Qwen 3.5 FP8 via vLLM.&lt;br /&gt;
* [https://github.com/ynotopec/api-rag '''api-rag'''] : API RAG (recherche sémantique) — fix support FAISS SVE, suppression logs verbeux, wait-for-APIS avant build index vectoriel.&lt;br /&gt;
* [https://github.com/ynotopec/api-diarization '''api-diarization'''] : API diarisation audio — install améliorée, ergonomics auth, documentation setup projet.&lt;br /&gt;
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] : API FastAPI ASR ( Whisper ) — refonte complète, installer uv, docs runtime.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-openmed '''privacy-proxy-openmed'''] : Privacy proxy OpenMed-based — LLM upstream optionnel (mode privacy-only).&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''privacy-proxy-gliner2'''] : Privacy proxy GLiNER2-based — détection entités nommées, upstream LLM optionnel, PR #1 en cours.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] : Proxy LLM custom — paramètres modèle configurables, alias ai-default, gestion erreurs upstream.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/block/buzz '''buzz (Block)'''] - Plateforme de communication &amp;quot;hive mind&amp;quot; : humains et agents IA cohabitent dans les memes canaux, auto-hebergeable.&lt;br /&gt;
* [https://github.com/pingdotgg/t3code '''t3code'''] - Interface web GUI pour orchestrer les agents de codage (Codex, Claude Code, Cursor...) depuis un navigateur.&lt;br /&gt;
* [https://github.com/moeru-ai/airi '''airi (Moeru AI)'''] - Companion IA auto-heberge : voice chat temps reel, controle de jeux (Minecraft, Factorio), compatible Grok.&lt;br /&gt;
* [https://github.com/bojieli/ai-agent-book '''ai-agent-book'''] - Livre open-source (10 chapitres, 94 experiences) couvrant les AI Agents : LLM + contexte + outils, de la theorie a la production.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd''' ] - Skill pour agent de codage qui arrete le discour de cacher la reponse. Sortie ADHD-friendly.&lt;br /&gt;
* [https://github.com/agegr/pi-web '''pi-web''' ] - Interface web pour agent de codage PI avec commandes et interfaces.&lt;br /&gt;
* [https://github.com/mattpocock/skills '''skills''' ] - Skills provenant directement du repertoire .agents d'un ingenieur reelement experimente.&lt;br /&gt;
* [https://github.com/earthtojake/text-to-cad '''text-to-cad''' ] - Collection de skills agent pour CAO, robotique et conception hardware.&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] - Transforme tout PDF technique en competence Claude Code : pret a etudier, referencer et utiliser pendant le travail.. (13726 stars)&lt;br /&gt;
* [https://github.com/UditAkhourii/adhd '''adhd'''] - Skill agent de codage : tree-of-thought avec pruning, base sur Claude &amp;amp; Codex Agent SDK. Ideation divergente.. (2879 stars)&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/different-ai/openwork openwork] - Alternative open-source a Claude Cowork (base sur opencode) : harness d agents de codage auto-heberge, multi-agents et multi-modeles.* [https://github.com/citrolabs/ego-lite '''ego-lite'''] - Navigateur optimise pour les agents IA : automation avec etat de session, compatible Codex et Claude Code.. (6536 stars)&lt;br /&gt;
* [https://github.com/zhaoxuya520/reverse-skill '''reverse-skill''' ] - Pack de compétences IA avec routage intelligent pour reverse engineering, pentesting autorisé et recherche de sécurité.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/TencentCloud/TencentDB-Agent-Memory '''TencentDB-Agent-Memory''' ] — Hub de mémoire d'équipe pour agents IA : transforme conversations, docs et code en 4 actifs réutilisables (Chat Memory, Skill, LLM-Wiki, Code-Graph).&lt;br /&gt;
&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Superto''']&lt;br /&gt;
* [https://github.com/microsoft/VibeVoice '''VibeVoice (Microsoft)'''] - Framework vocal open-source (TTS + ASR) de Microsoft : synthese, clonage et reconnaissance vocale de '''pointe.nic-3''' — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/andrewyng/aisuite '''aisuite'''] - Interface unifiee pour les LLM : un seul appel OpenAI-compatible pour OpenAI, Anthropic, Google, AWS, Cohere, Ollama, OpenRouter...&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les L&lt;br /&gt;
* [https://github.com/opengeos/GeoLibre GeoLibre] - Plateforme SIG open-source cloud-native pour visualisation et analyse de donnees geospatiales (web, bureau, mobile, Jupyter).&lt;br /&gt;
* [https://github.com/CoreBunch/Instatic Instatic] - CMS visuel auto-heberge et agentic alternative a Webflow/Framer : utilisateurs, roles, plugins, base de donnees integree.LM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
* [https://github.com/lyogavin/airllm '''airllm''' ] - Inférence LLM haute performance sur un seul GPU 4Go, permettant l'exécution de modèles de 70B+ paramètres en local.&lt;br /&gt;
* [https://github.com/oblien/openship '''openship''' ] - Plateforme de deploiement auto-hebergeee et auto-deployee avec travaux automatiques.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://github.com/KnockOutEZ/wigolo '''wigolo''' ] - Web local-first pour agent de codage — recherche, fetch, crawl et recherche via MCP, sans clee API.&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2078</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2078"/>
		<updated>2026-08-04T22:01:41Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Ajout TencentDB-Agent-Memory (trending hebdo 2026-08-04) - Hub mémoire d'équipe pour agents IA&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com/docs '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 02/06 - 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] : API de traduction multilingue temps réel — PR #3 fix timeout (gestion réponses lentes), PR #2 URLs configurables via .env.&lt;br /&gt;
* [https://github.com/ynotopec/GLM-4.7-Flash-vLLM '''GLM-4.7-Flash-vLLM'''] : Déploiement vLLM GLM-4.7 Flash.&lt;br /&gt;
* [https://github.com/ynotopec/qwen3.5-fp8-vllm '''qwen3.5-fp8-vllm'''] : Déploiement Qwen 3.5 FP8 via vLLM.&lt;br /&gt;
* [https://github.com/ynotopec/api-rag '''api-rag'''] : API RAG (recherche sémantique) — fix support FAISS SVE, suppression logs verbeux, wait-for-APIS avant build index vectoriel.&lt;br /&gt;
* [https://github.com/ynotopec/api-diarization '''api-diarization'''] : API diarisation audio — install améliorée, ergonomics auth, documentation setup projet.&lt;br /&gt;
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] : API FastAPI ASR ( Whisper ) — refonte complète, installer uv, docs runtime.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-openmed '''privacy-proxy-openmed'''] : Privacy proxy OpenMed-based — LLM upstream optionnel (mode privacy-only).&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''privacy-proxy-gliner2'''] : Privacy proxy GLiNER2-based — détection entités nommées, upstream LLM optionnel, PR #1 en cours.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] : Proxy LLM custom — paramètres modèle configurables, alias ai-default, gestion erreurs upstream.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/block/buzz '''buzz (Block)'''] - Plateforme de communication &amp;quot;hive mind&amp;quot; : humains et agents IA cohabitent dans les memes canaux, auto-hebergeable.&lt;br /&gt;
* [https://github.com/pingdotgg/t3code '''t3code'''] - Interface web GUI pour orchestrer les agents de codage (Codex, Claude Code, Cursor...) depuis un navigateur.&lt;br /&gt;
* [https://github.com/moeru-ai/airi '''airi (Moeru AI)'''] - Companion IA auto-heberge : voice chat temps reel, controle de jeux (Minecraft, Factorio), compatible Grok.&lt;br /&gt;
* [https://github.com/bojieli/ai-agent-book '''ai-agent-book'''] - Livre open-source (10 chapitres, 94 experiences) couvrant les AI Agents : LLM + contexte + outils, de la theorie a la production.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd''' ] - Skill pour agent de codage qui arrete le discour de cacher la reponse. Sortie ADHD-friendly.&lt;br /&gt;
* [https://github.com/agegr/pi-web '''pi-web''' ] - Interface web pour agent de codage PI avec commandes et interfaces.&lt;br /&gt;
* [https://github.com/mattpocock/skills '''skills''' ] - Skills provenant directement du repertoire .agents d'un ingenieur reelement experimente.&lt;br /&gt;
* [https://github.com/earthtojake/text-to-cad '''text-to-cad''' ] - Collection de skills agent pour CAO, robotique et conception hardware.&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] - Transforme tout PDF technique en competence Claude Code : pret a etudier, referencer et utiliser pendant le travail.. (13726 stars)&lt;br /&gt;
* [https://github.com/UditAkhourii/adhd '''adhd'''] - Skill agent de codage : tree-of-thought avec pruning, base sur Claude &amp;amp; Codex Agent SDK. Ideation divergente.. (2879 stars)&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/different-ai/openwork openwork] - Alternative open-source a Claude Cowork (base sur opencode) : harness d agents de codage auto-heberge, multi-agents et multi-modeles.* [https://github.com/citrolabs/ego-lite '''ego-lite'''] - Navigateur optimise pour les agents IA : automation avec etat de session, compatible Codex et Claude Code.. (6536 stars)&lt;br /&gt;
* [https://github.com/zhaoxuya520/reverse-skill '''reverse-skill''' ] - Pack de compétences IA avec routage intelligent pour reverse engineering, pentesting autorisé et recherche de sécurité.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/TencentCloud/TencentDB-Agent-Memory '''TencentDB-Agent-Memory''' ] — Hub de mémoire d'équipe pour agents IA : transforme conversations, docs et code en 4 actifs réutilisables (Chat Memory, Skill, LLM-Wiki, Code-Graph).&lt;br /&gt;
&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Superto''']&lt;br /&gt;
* [https://github.com/microsoft/VibeVoice '''VibeVoice (Microsoft)'''] - Framework vocal open-source (TTS + ASR) de Microsoft : synthese, clonage et reconnaissance vocale de '''pointe.nic-3''' — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/andrewyng/aisuite '''aisuite'''] - Interface unifiee pour les LLM : un seul appel OpenAI-compatible pour OpenAI, Anthropic, Google, AWS, Cohere, Ollama, OpenRouter...&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les L&lt;br /&gt;
* [https://github.com/opengeos/GeoLibre GeoLibre] - Plateforme SIG open-source cloud-native pour visualisation et analyse de donnees geospatiales (web, bureau, mobile, Jupyter).&lt;br /&gt;
* [https://github.com/CoreBunch/Instatic Instatic] - CMS visuel auto-heberge et agentic alternative a Webflow/Framer : utilisateurs, roles, plugins, base de donnees integree.LM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
* [https://github.com/lyogavin/airllm '''airllm''' ] - Inférence LLM haute performance sur un seul GPU 4Go, permettant l'exécution de modèles de 70B+ paramètres en local.&lt;br /&gt;
* [https://github.com/oblien/openship '''openship''' ] - Plateforme de deploiement auto-hebergeee et auto-deployee avec travaux automatiques.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://github.com/KnockOutEZ/wigolo '''wigolo''' ] - Web local-first pour agent de codage — recherche, fetch, crawl et recherche via MCP, sans clee API.&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6 fp8''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6 fp8''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2077</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2077"/>
		<updated>2026-08-04T22:01:04Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Ajout TencentDB-Agent-Memory (trending hebdo) - Hub mémoire d'équipe pour agents IA&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;@/tmp/wiki_main_edited.txt&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2076</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2076"/>
		<updated>2026-08-04T01:28:32Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Ajout projets trending GitHub 04/08 - Agents IA et Infrastructure&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com/docs '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 02/06 - 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] : API de traduction multilingue temps réel — PR #3 fix timeout (gestion réponses lentes), PR #2 URLs configurables via .env.&lt;br /&gt;
* [https://github.com/ynotopec/GLM-4.7-Flash-vLLM '''GLM-4.7-Flash-vLLM'''] : Déploiement vLLM GLM-4.7 Flash.&lt;br /&gt;
* [https://github.com/ynotopec/qwen3.5-fp8-vllm '''qwen3.5-fp8-vllm'''] : Déploiement Qwen 3.5 FP8 via vLLM.&lt;br /&gt;
* [https://github.com/ynotopec/api-rag '''api-rag'''] : API RAG (recherche sémantique) — fix support FAISS SVE, suppression logs verbeux, wait-for-APIS avant build index vectoriel.&lt;br /&gt;
* [https://github.com/ynotopec/api-diarization '''api-diarization'''] : API diarisation audio — install améliorée, ergonomics auth, documentation setup projet.&lt;br /&gt;
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] : API FastAPI ASR ( Whisper ) — refonte complète, installer uv, docs runtime.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-openmed '''privacy-proxy-openmed'''] : Privacy proxy OpenMed-based — LLM upstream optionnel (mode privacy-only).&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''privacy-proxy-gliner2'''] : Privacy proxy GLiNER2-based — détection entités nommées, upstream LLM optionnel, PR #1 en cours.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] : Proxy LLM custom — paramètres modèle configurables, alias ai-default, gestion erreurs upstream.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/block/buzz '''buzz (Block)'''] - Plateforme de communication &amp;quot;hive mind&amp;quot; : humains et agents IA cohabitent dans les memes canaux, auto-hebergeable.&lt;br /&gt;
* [https://github.com/pingdotgg/t3code '''t3code'''] - Interface web GUI pour orchestrer les agents de codage (Codex, Claude Code, Cursor...) depuis un navigateur.&lt;br /&gt;
* [https://github.com/moeru-ai/airi '''airi (Moeru AI)'''] - Companion IA auto-heberge : voice chat temps reel, controle de jeux (Minecraft, Factorio), compatible Grok.&lt;br /&gt;
* [https://github.com/bojieli/ai-agent-book '''ai-agent-book'''] - Livre open-source (10 chapitres, 94 experiences) couvrant les AI Agents : LLM + contexte + outils, de la theorie a la production.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd''' ] - Skill pour agent de codage qui arrete le discour de cacher la reponse. Sortie ADHD-friendly.&lt;br /&gt;
* [https://github.com/agegr/pi-web '''pi-web''' ] - Interface web pour agent de codage PI avec commandes et interfaces.&lt;br /&gt;
* [https://github.com/mattpocock/skills '''skills''' ] - Skills provenant directement du repertoire .agents d'un ingenieur reelement experimente.&lt;br /&gt;
* [https://github.com/earthtojake/text-to-cad '''text-to-cad''' ] - Collection de skills agent pour CAO, robotique et conception hardware.&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] - Transforme tout PDF technique en competence Claude Code : pret a etudier, referencer et utiliser pendant le travail.. (13726 stars)&lt;br /&gt;
* [https://github.com/UditAkhourii/adhd '''adhd'''] - Skill agent de codage : tree-of-thought avec pruning, base sur Claude &amp;amp; Codex Agent SDK. Ideation divergente.. (2879 stars)&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/different-ai/openwork openwork] - Alternative open-source a Claude Cowork (base sur opencode) : harness d agents de codage auto-heberge, multi-agents et multi-modeles.* [https://github.com/citrolabs/ego-lite '''ego-lite'''] - Navigateur optimise pour les agents IA : automation avec etat de session, compatible Codex et Claude Code.. (6536 stars)&lt;br /&gt;
* [https://github.com/zhaoxuya520/reverse-skill '''reverse-skill''' ] - Pack de compétences IA avec routage intelligent pour reverse engineering, pentesting autorisé et recherche de sécurité.&lt;br /&gt;
&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Superto''']&lt;br /&gt;
* [https://github.com/microsoft/VibeVoice '''VibeVoice (Microsoft)'''] - Framework vocal open-source (TTS + ASR) de Microsoft : synthese, clonage et reconnaissance vocale de '''pointe.nic-3''' — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/andrewyng/aisuite '''aisuite'''] - Interface unifiee pour les LLM : un seul appel OpenAI-compatible pour OpenAI, Anthropic, Google, AWS, Cohere, Ollama, OpenRouter...&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les L&lt;br /&gt;
* [https://github.com/opengeos/GeoLibre GeoLibre] - Plateforme SIG open-source cloud-native pour visualisation et analyse de donnees geospatiales (web, bureau, mobile, Jupyter).&lt;br /&gt;
* [https://github.com/CoreBunch/Instatic Instatic] - CMS visuel auto-heberge et agentic alternative a Webflow/Framer : utilisateurs, roles, plugins, base de donnees integree.LM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
* [https://github.com/lyogavin/airllm '''airllm''' ] - Inférence LLM haute performance sur un seul GPU 4Go, permettant l'exécution de modèles de 70B+ paramètres en local.&lt;br /&gt;
* [https://github.com/oblien/openship '''openship''' ] - Plateforme de deploiement auto-hebergeee et auto-deployee avec travaux automatiques.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://github.com/KnockOutEZ/wigolo '''wigolo''' ] - Web local-first pour agent de codage — recherche, fetch, crawl et recherche via MCP, sans clee API.&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6 fp8''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6 fp8''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2075</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2075"/>
		<updated>2026-08-03T12:18:01Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: /* Nouveautés 01/08/2026 */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com/docs '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 02/06 - 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] : API de traduction multilingue temps réel — PR #3 fix timeout (gestion réponses lentes), PR #2 URLs configurables via .env.&lt;br /&gt;
* [https://github.com/ynotopec/GLM-4.7-Flash-vLLM '''GLM-4.7-Flash-vLLM'''] : Déploiement vLLM GLM-4.7 Flash.&lt;br /&gt;
* [https://github.com/ynotopec/qwen3.5-fp8-vllm '''qwen3.5-fp8-vllm'''] : Déploiement Qwen 3.5 FP8 via vLLM.&lt;br /&gt;
* [https://github.com/ynotopec/api-rag '''api-rag'''] : API RAG (recherche sémantique) — fix support FAISS SVE, suppression logs verbeux, wait-for-APIS avant build index vectoriel.&lt;br /&gt;
* [https://github.com/ynotopec/api-diarization '''api-diarization'''] : API diarisation audio — install améliorée, ergonomics auth, documentation setup projet.&lt;br /&gt;
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] : API FastAPI ASR ( Whisper ) — refonte complète, installer uv, docs runtime.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-openmed '''privacy-proxy-openmed'''] : Privacy proxy OpenMed-based — LLM upstream optionnel (mode privacy-only).&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''privacy-proxy-gliner2'''] : Privacy proxy GLiNER2-based — détection entités nommées, upstream LLM optionnel, PR #1 en cours.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] : Proxy LLM custom — paramètres modèle configurables, alias ai-default, gestion erreurs upstream.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/block/buzz '''buzz (Block)'''] - Plateforme de communication &amp;quot;hive mind&amp;quot; : humains et agents IA cohabitent dans les memes canaux, auto-hebergeable.&lt;br /&gt;
* [https://github.com/pingdotgg/t3code '''t3code'''] - Interface web GUI pour orchestrer les agents de codage (Codex, Claude Code, Cursor...) depuis un navigateur.&lt;br /&gt;
* [https://github.com/moeru-ai/airi '''airi (Moeru AI)'''] - Companion IA auto-heberge : voice chat temps reel, controle de jeux (Minecraft, Factorio), compatible Grok.&lt;br /&gt;
* [https://github.com/bojieli/ai-agent-book '''ai-agent-book'''] - Livre open-source (10 chapitres, 94 experiences) couvrant les AI Agents : LLM + contexte + outils, de la theorie a la production.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd''' ] - Skill pour agent de codage qui arrete le discour de cacher la reponse. Sortie ADHD-friendly.&lt;br /&gt;
* [https://github.com/agegr/pi-web '''pi-web''' ] - Interface web pour agent de codage PI avec commandes et interfaces.&lt;br /&gt;
* [https://github.com/mattpocock/skills '''skills''' ] - Skills provenant directement du repertoire .agents d'un ingenieur reelement experimente.&lt;br /&gt;
* [https://github.com/earthtojake/text-to-cad '''text-to-cad''' ] - Collection de skills agent pour CAO, robotique et conception hardware.&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] - Transforme tout PDF technique en competence Claude Code : pret a etudier, referencer et utiliser pendant le travail.. (13726 stars)&lt;br /&gt;
* [https://github.com/UditAkhourii/adhd '''adhd'''] - Skill agent de codage : tree-of-thought avec pruning, base sur Claude &amp;amp; Codex Agent SDK. Ideation divergente.. (2879 stars)&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/different-ai/openwork openwork] - Alternative open-source a Claude Cowork (base sur opencode) : harness d agents de codage auto-heberge, multi-agents et multi-modeles.* [https://github.com/citrolabs/ego-lite '''ego-lite'''] - Navigateur optimise pour les agents IA : automation avec etat de session, compatible Codex et Claude Code.. (6536 stars)&lt;br /&gt;
&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Superto''']&lt;br /&gt;
* [https://github.com/microsoft/VibeVoice '''VibeVoice (Microsoft)'''] - Framework vocal open-source (TTS + ASR) de Microsoft : synthese, clonage et reconnaissance vocale de '''pointe.nic-3''' — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/andrewyng/aisuite '''aisuite'''] - Interface unifiee pour les LLM : un seul appel OpenAI-compatible pour OpenAI, Anthropic, Google, AWS, Cohere, Ollama, OpenRouter...&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les L&lt;br /&gt;
* [https://github.com/opengeos/GeoLibre GeoLibre] - Plateforme SIG open-source cloud-native pour visualisation et analyse de donnees geospatiales (web, bureau, mobile, Jupyter).&lt;br /&gt;
* [https://github.com/CoreBunch/Instatic Instatic] - CMS visuel auto-heberge et agentic alternative a Webflow/Framer : utilisateurs, roles, plugins, base de donnees integree.LM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
* [https://github.com/oblien/openship '''openship''' ] - Plateforme de deploiement auto-hebergeee et auto-deployee avec travaux automatiques.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://github.com/KnockOutEZ/wigolo '''wigolo''' ] - Web local-first pour agent de codage — recherche, fetch, crawl et recherche via MCP, sans clee API.&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6 fp8''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6 fp8''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2074</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2074"/>
		<updated>2026-08-03T11:18:50Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: /* Audio &amp;amp; TTS */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 02/06 - 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] : API de traduction multilingue temps réel — PR #3 fix timeout (gestion réponses lentes), PR #2 URLs configurables via .env.&lt;br /&gt;
* [https://github.com/ynotopec/GLM-4.7-Flash-vLLM '''GLM-4.7-Flash-vLLM'''] : Déploiement vLLM GLM-4.7 Flash.&lt;br /&gt;
* [https://github.com/ynotopec/qwen3.5-fp8-vllm '''qwen3.5-fp8-vllm'''] : Déploiement Qwen 3.5 FP8 via vLLM.&lt;br /&gt;
* [https://github.com/ynotopec/api-rag '''api-rag'''] : API RAG (recherche sémantique) — fix support FAISS SVE, suppression logs verbeux, wait-for-APIS avant build index vectoriel.&lt;br /&gt;
* [https://github.com/ynotopec/api-diarization '''api-diarization'''] : API diarisation audio — install améliorée, ergonomics auth, documentation setup projet.&lt;br /&gt;
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] : API FastAPI ASR ( Whisper ) — refonte complète, installer uv, docs runtime.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-openmed '''privacy-proxy-openmed'''] : Privacy proxy OpenMed-based — LLM upstream optionnel (mode privacy-only).&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''privacy-proxy-gliner2'''] : Privacy proxy GLiNER2-based — détection entités nommées, upstream LLM optionnel, PR #1 en cours.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] : Proxy LLM custom — paramètres modèle configurables, alias ai-default, gestion erreurs upstream.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/block/buzz '''buzz (Block)'''] - Plateforme de communication &amp;quot;hive mind&amp;quot; : humains et agents IA cohabitent dans les memes canaux, auto-hebergeable.&lt;br /&gt;
* [https://github.com/pingdotgg/t3code '''t3code'''] - Interface web GUI pour orchestrer les agents de codage (Codex, Claude Code, Cursor...) depuis un navigateur.&lt;br /&gt;
* [https://github.com/moeru-ai/airi '''airi (Moeru AI)'''] - Companion IA auto-heberge : voice chat temps reel, controle de jeux (Minecraft, Factorio), compatible Grok.&lt;br /&gt;
* [https://github.com/bojieli/ai-agent-book '''ai-agent-book'''] - Livre open-source (10 chapitres, 94 experiences) couvrant les AI Agents : LLM + contexte + outils, de la theorie a la production.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd''' ] - Skill pour agent de codage qui arrete le discour de cacher la reponse. Sortie ADHD-friendly.&lt;br /&gt;
* [https://github.com/agegr/pi-web '''pi-web''' ] - Interface web pour agent de codage PI avec commandes et interfaces.&lt;br /&gt;
* [https://github.com/mattpocock/skills '''skills''' ] - Skills provenant directement du repertoire .agents d'un ingenieur reelement experimente.&lt;br /&gt;
* [https://github.com/earthtojake/text-to-cad '''text-to-cad''' ] - Collection de skills agent pour CAO, robotique et conception hardware.&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] - Transforme tout PDF technique en competence Claude Code : pret a etudier, referencer et utiliser pendant le travail.. (13726 stars)&lt;br /&gt;
* [https://github.com/UditAkhourii/adhd '''adhd'''] - Skill agent de codage : tree-of-thought avec pruning, base sur Claude &amp;amp; Codex Agent SDK. Ideation divergente.. (2879 stars)&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/different-ai/openwork openwork] - Alternative open-source a Claude Cowork (base sur opencode) : harness d agents de codage auto-heberge, multi-agents et multi-modeles.* [https://github.com/citrolabs/ego-lite '''ego-lite'''] - Navigateur optimise pour les agents IA : automation avec etat de session, compatible Codex et Claude Code.. (6536 stars)&lt;br /&gt;
&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Superto''']&lt;br /&gt;
* [https://github.com/microsoft/VibeVoice '''VibeVoice (Microsoft)'''] - Framework vocal open-source (TTS + ASR) de Microsoft : synthese, clonage et reconnaissance vocale de '''pointe.nic-3''' — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/andrewyng/aisuite '''aisuite'''] - Interface unifiee pour les LLM : un seul appel OpenAI-compatible pour OpenAI, Anthropic, Google, AWS, Cohere, Ollama, OpenRouter...&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les L&lt;br /&gt;
* [https://github.com/opengeos/GeoLibre GeoLibre] - Plateforme SIG open-source cloud-native pour visualisation et analyse de donnees geospatiales (web, bureau, mobile, Jupyter).&lt;br /&gt;
* [https://github.com/CoreBunch/Instatic Instatic] - CMS visuel auto-heberge et agentic alternative a Webflow/Framer : utilisateurs, roles, plugins, base de donnees integree.LM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
* [https://github.com/oblien/openship '''openship''' ] - Plateforme de deploiement auto-hebergeee et auto-deployee avec travaux automatiques.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://github.com/KnockOutEZ/wigolo '''wigolo''' ] - Web local-first pour agent de codage — recherche, fetch, crawl et recherche via MCP, sans clee API.&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6 fp8''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6 fp8''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2073</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2073"/>
		<updated>2026-08-03T09:39:01Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: /* Audio &amp;amp; TTS */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 02/06 - 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] : API de traduction multilingue temps réel — PR #3 fix timeout (gestion réponses lentes), PR #2 URLs configurables via .env.&lt;br /&gt;
* [https://github.com/ynotopec/GLM-4.7-Flash-vLLM '''GLM-4.7-Flash-vLLM'''] : Déploiement vLLM GLM-4.7 Flash.&lt;br /&gt;
* [https://github.com/ynotopec/qwen3.5-fp8-vllm '''qwen3.5-fp8-vllm'''] : Déploiement Qwen 3.5 FP8 via vLLM.&lt;br /&gt;
* [https://github.com/ynotopec/api-rag '''api-rag'''] : API RAG (recherche sémantique) — fix support FAISS SVE, suppression logs verbeux, wait-for-APIS avant build index vectoriel.&lt;br /&gt;
* [https://github.com/ynotopec/api-diarization '''api-diarization'''] : API diarisation audio — install améliorée, ergonomics auth, documentation setup projet.&lt;br /&gt;
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] : API FastAPI ASR ( Whisper ) — refonte complète, installer uv, docs runtime.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-openmed '''privacy-proxy-openmed'''] : Privacy proxy OpenMed-based — LLM upstream optionnel (mode privacy-only).&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''privacy-proxy-gliner2'''] : Privacy proxy GLiNER2-based — détection entités nommées, upstream LLM optionnel, PR #1 en cours.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] : Proxy LLM custom — paramètres modèle configurables, alias ai-default, gestion erreurs upstream.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/block/buzz '''buzz (Block)'''] - Plateforme de communication &amp;quot;hive mind&amp;quot; : humains et agents IA cohabitent dans les memes canaux, auto-hebergeable.&lt;br /&gt;
* [https://github.com/pingdotgg/t3code '''t3code'''] - Interface web GUI pour orchestrer les agents de codage (Codex, Claude Code, Cursor...) depuis un navigateur.&lt;br /&gt;
* [https://github.com/moeru-ai/airi '''airi (Moeru AI)'''] - Companion IA auto-heberge : voice chat temps reel, controle de jeux (Minecraft, Factorio), compatible Grok.&lt;br /&gt;
* [https://github.com/bojieli/ai-agent-book '''ai-agent-book'''] - Livre open-source (10 chapitres, 94 experiences) couvrant les AI Agents : LLM + contexte + outils, de la theorie a la production.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd''' ] - Skill pour agent de codage qui arrete le discour de cacher la reponse. Sortie ADHD-friendly.&lt;br /&gt;
* [https://github.com/agegr/pi-web '''pi-web''' ] - Interface web pour agent de codage PI avec commandes et interfaces.&lt;br /&gt;
* [https://github.com/mattpocock/skills '''skills''' ] - Skills provenant directement du repertoire .agents d'un ingenieur reelement experimente.&lt;br /&gt;
* [https://github.com/earthtojake/text-to-cad '''text-to-cad''' ] - Collection de skills agent pour CAO, robotique et conception hardware.&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] - Transforme tout PDF technique en competence Claude Code : pret a etudier, referencer et utiliser pendant le travail.. (13726 stars)&lt;br /&gt;
* [https://github.com/UditAkhourii/adhd '''adhd'''] - Skill agent de codage : tree-of-thought avec pruning, base sur Claude &amp;amp; Codex Agent SDK. Ideation divergente.. (2879 stars)&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/different-ai/openwork openwork] - Alternative open-source a Claude Cowork (base sur opencode) : harness d agents de codage auto-heberge, multi-agents et multi-modeles.* [https://github.com/citrolabs/ego-lite '''ego-lite'''] - Navigateur optimise pour les agents IA : automation avec etat de session, compatible Codex et Claude Code.. (6536 stars)&lt;br /&gt;
&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Superto''']&lt;br /&gt;
* [https://github.com/microsoft/VibeVoice VibeVoice (Microsoft)] - Framework vocal open-source (TTS + ASR) de Microsoft : synthese, clonage et reconnaissance vocale de pointe.nic-3'''] — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/andrewyng/aisuite '''aisuite'''] - Interface unifiee pour les LLM : un seul appel OpenAI-compatible pour OpenAI, Anthropic, Google, AWS, Cohere, Ollama, OpenRouter...&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les L&lt;br /&gt;
* [https://github.com/opengeos/GeoLibre GeoLibre] - Plateforme SIG open-source cloud-native pour visualisation et analyse de donnees geospatiales (web, bureau, mobile, Jupyter).&lt;br /&gt;
* [https://github.com/CoreBunch/Instatic Instatic] - CMS visuel auto-heberge et agentic alternative a Webflow/Framer : utilisateurs, roles, plugins, base de donnees integree.LM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
* [https://github.com/oblien/openship '''openship''' ] - Plateforme de deploiement auto-hebergeee et auto-deployee avec travaux automatiques.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://github.com/KnockOutEZ/wigolo '''wigolo''' ] - Web local-first pour agent de codage — recherche, fetch, crawl et recherche via MCP, sans clee API.&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6 fp8''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6 fp8''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2072</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2072"/>
		<updated>2026-08-03T01:28:00Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Update: ajout 5 nouveaux projets trending GitHub - 03/08/2026&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 02/06 - 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] : API de traduction multilingue temps réel — PR #3 fix timeout (gestion réponses lentes), PR #2 URLs configurables via .env.&lt;br /&gt;
* [https://github.com/ynotopec/GLM-4.7-Flash-vLLM '''GLM-4.7-Flash-vLLM'''] : Déploiement vLLM GLM-4.7 Flash.&lt;br /&gt;
* [https://github.com/ynotopec/qwen3.5-fp8-vllm '''qwen3.5-fp8-vllm'''] : Déploiement Qwen 3.5 FP8 via vLLM.&lt;br /&gt;
* [https://github.com/ynotopec/api-rag '''api-rag'''] : API RAG (recherche sémantique) — fix support FAISS SVE, suppression logs verbeux, wait-for-APIS avant build index vectoriel.&lt;br /&gt;
* [https://github.com/ynotopec/api-diarization '''api-diarization'''] : API diarisation audio — install améliorée, ergonomics auth, documentation setup projet.&lt;br /&gt;
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] : API FastAPI ASR ( Whisper ) — refonte complète, installer uv, docs runtime.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-openmed '''privacy-proxy-openmed'''] : Privacy proxy OpenMed-based — LLM upstream optionnel (mode privacy-only).&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''privacy-proxy-gliner2'''] : Privacy proxy GLiNER2-based — détection entités nommées, upstream LLM optionnel, PR #1 en cours.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] : Proxy LLM custom — paramètres modèle configurables, alias ai-default, gestion erreurs upstream.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/block/buzz '''buzz (Block)'''] - Plateforme de communication &amp;quot;hive mind&amp;quot; : humains et agents IA cohabitent dans les memes canaux, auto-hebergeable.&lt;br /&gt;
* [https://github.com/pingdotgg/t3code '''t3code'''] - Interface web GUI pour orchestrer les agents de codage (Codex, Claude Code, Cursor...) depuis un navigateur.&lt;br /&gt;
* [https://github.com/moeru-ai/airi '''airi (Moeru AI)'''] - Companion IA auto-heberge : voice chat temps reel, controle de jeux (Minecraft, Factorio), compatible Grok.&lt;br /&gt;
* [https://github.com/bojieli/ai-agent-book '''ai-agent-book'''] - Livre open-source (10 chapitres, 94 experiences) couvrant les AI Agents : LLM + contexte + outils, de la theorie a la production.&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd''' ] - Skill pour agent de codage qui arrete le discour de cacher la reponse. Sortie ADHD-friendly.&lt;br /&gt;
* [https://github.com/agegr/pi-web '''pi-web''' ] - Interface web pour agent de codage PI avec commandes et interfaces.&lt;br /&gt;
* [https://github.com/mattpocock/skills '''skills''' ] - Skills provenant directement du repertoire .agents d'un ingenieur reelement experimente.&lt;br /&gt;
* [https://github.com/earthtojake/text-to-cad '''text-to-cad''' ] - Collection de skills agent pour CAO, robotique et conception hardware.&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] - Transforme tout PDF technique en competence Claude Code : pret a etudier, referencer et utiliser pendant le travail.. (13726 stars)&lt;br /&gt;
* [https://github.com/UditAkhourii/adhd '''adhd'''] - Skill agent de codage : tree-of-thought avec pruning, base sur Claude &amp;amp; Codex Agent SDK. Ideation divergente.. (2879 stars)&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/different-ai/openwork openwork] - Alternative open-source a Claude Cowork (base sur opencode) : harness d agents de codage auto-heberge, multi-agents et multi-modeles.* [https://github.com/citrolabs/ego-lite '''ego-lite'''] - Navigateur optimise pour les agents IA : automation avec etat de session, compatible Codex et Claude Code.. (6536 stars)&lt;br /&gt;
&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Superto&lt;br /&gt;
* [https://github.com/microsoft/VibeVoice VibeVoice (Microsoft)] - Framework vocal open-source (TTS + ASR) de Microsoft : synthese, clonage et reconnaissance vocale de pointe.nic-3'''] — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/andrewyng/aisuite '''aisuite'''] - Interface unifiee pour les LLM : un seul appel OpenAI-compatible pour OpenAI, Anthropic, Google, AWS, Cohere, Ollama, OpenRouter...&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les L&lt;br /&gt;
* [https://github.com/opengeos/GeoLibre GeoLibre] - Plateforme SIG open-source cloud-native pour visualisation et analyse de donnees geospatiales (web, bureau, mobile, Jupyter).&lt;br /&gt;
* [https://github.com/CoreBunch/Instatic Instatic] - CMS visuel auto-heberge et agentic alternative a Webflow/Framer : utilisateurs, roles, plugins, base de donnees integree.LM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
* [https://github.com/oblien/openship '''openship''' ] - Plateforme de deploiement auto-hebergeee et auto-deployee avec travaux automatiques.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://github.com/KnockOutEZ/wigolo '''wigolo''' ] - Web local-first pour agent de codage — recherche, fetch, crawl et recherche via MCP, sans clee API.&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6 fp8''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6 fp8''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2071</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2071"/>
		<updated>2026-08-02T06:24:11Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Ajout: 9 nouveaux repos GitHub (translate-rt, GLM-4.7, qwen3.5-fp8, api-rag, api-diarization, api-audio2txt, api-llm-privacy-proxy-openmed, api-llm-privacy-proxy-gliner2, api-llm-custom) — mise à jour 2 mois&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 02/06 - 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] : API de traduction multilingue temps réel — PR #3 fix timeout (gestion réponses lentes), PR #2 URLs configurables via .env.&lt;br /&gt;
* [https://github.com/ynotopec/GLM-4.7-Flash-vLLM '''GLM-4.7-Flash-vLLM'''] : Déploiement vLLM GLM-4.7 Flash.&lt;br /&gt;
* [https://github.com/ynotopec/qwen3.5-fp8-vllm '''qwen3.5-fp8-vllm'''] : Déploiement Qwen 3.5 FP8 via vLLM.&lt;br /&gt;
* [https://github.com/ynotopec/api-rag '''api-rag'''] : API RAG (recherche sémantique) — fix support FAISS SVE, suppression logs verbeux, wait-for-APIS avant build index vectoriel.&lt;br /&gt;
* [https://github.com/ynotopec/api-diarization '''api-diarization'''] : API diarisation audio — install améliorée, ergonomics auth, documentation setup projet.&lt;br /&gt;
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] : API FastAPI ASR ( Whisper ) — refonte complète, installer uv, docs runtime.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-openmed '''privacy-proxy-openmed'''] : Privacy proxy OpenMed-based — LLM upstream optionnel (mode privacy-only).&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''privacy-proxy-gliner2'''] : Privacy proxy GLiNER2-based — détection entités nommées, upstream LLM optionnel, PR #1 en cours.&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] : Proxy LLM custom — paramètres modèle configurables, alias ai-default, gestion erreurs upstream.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd''' ] - Skill pour agent de codage qui arrete le discour de cacher la reponse. Sortie ADHD-friendly.&lt;br /&gt;
* [https://github.com/agegr/pi-web '''pi-web''' ] - Interface web pour agent de codage PI avec commandes et interfaces.&lt;br /&gt;
* [https://github.com/mattpocock/skills '''skills''' ] - Skills provenant directement du repertoire .agents d'un ingenieur reelement experimente.&lt;br /&gt;
* [https://github.com/earthtojake/text-to-cad '''text-to-cad''' ] - Collection de skills agent pour CAO, robotique et conception hardware.&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] - Transforme tout PDF technique en competence Claude Code : pret a etudier, referencer et utiliser pendant le travail.. (13726 stars)&lt;br /&gt;
* [https://github.com/UditAkhourii/adhd '''adhd'''] - Skill agent de codage : tree-of-thought avec pruning, base sur Claude &amp;amp; Codex Agent SDK. Ideation divergente.. (2879 stars)&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/different-ai/openwork openwork] - Alternative open-source a Claude Cowork (base sur opencode) : harness d agents de codage auto-heberge, multi-agents et multi-modeles.* [https://github.com/citrolabs/ego-lite '''ego-lite'''] - Navigateur optimise pour les agents IA : automation avec etat de session, compatible Codex et Claude Code.. (6536 stars)&lt;br /&gt;
&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Superto&lt;br /&gt;
* [https://github.com/microsoft/VibeVoice VibeVoice (Microsoft)] - Framework vocal open-source (TTS + ASR) de Microsoft : synthese, clonage et reconnaissance vocale de pointe.nic-3'''] — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les L&lt;br /&gt;
* [https://github.com/opengeos/GeoLibre GeoLibre] - Plateforme SIG open-source cloud-native pour visualisation et analyse de donnees geospatiales (web, bureau, mobile, Jupyter).&lt;br /&gt;
* [https://github.com/CoreBunch/Instatic Instatic] - CMS visuel auto-heberge et agentic alternative a Webflow/Framer : utilisateurs, roles, plugins, base de donnees integree.LM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
* [https://github.com/oblien/openship '''openship''' ] - Plateforme de deploiement auto-hebergeee et auto-deployee avec travaux automatiques.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://github.com/KnockOutEZ/wigolo '''wigolo''' ] - Web local-first pour agent de codage — recherche, fetch, crawl et recherche via MCP, sans clee API.&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6 fp8''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6 fp8''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2070</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2070"/>
		<updated>2026-08-02T06:14:37Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Update: ajouter infocepo-infra-mcp, nemotron-embed, agent-saas (commits récents GitHub)&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 03/08/2026 ==&lt;br /&gt;
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] : MCP server universel — zéro clé en dur, lecture API key depuis ~/.infocepo-credentials, déploiement K8s, refacto stdio (plus de shell wrapper).&lt;br /&gt;
* [https://github.com/ynotopec/nemotron-embed '''nemotron-embed'''] : Service Nemotron-3 Embedding vLLM — systemd service, script de lancement, documentation d'installation.&lt;br /&gt;
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] : Plateforme SaaS d'agents IA autonomes — manifests K8s complets, dashboard FastAPI, hardening déploiement.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd''' ] - Skill pour agent de codage qui arrete le discour de cacher la reponse. Sortie ADHD-friendly.&lt;br /&gt;
* [https://github.com/agegr/pi-web '''pi-web''' ] - Interface web pour agent de codage PI avec commandes et interfaces.&lt;br /&gt;
* [https://github.com/mattpocock/skills '''skills''' ] - Skills provenant directement du repertoire .agents d'un ingenieur reelement experimente.&lt;br /&gt;
* [https://github.com/earthtojake/text-to-cad '''text-to-cad''' ] - Collection de skills agent pour CAO, robotique et conception hardware.&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] - Transforme tout PDF technique en competence Claude Code : pret a etudier, referencer et utiliser pendant le travail.. (13726 stars)&lt;br /&gt;
* [https://github.com/UditAkhourii/adhd '''adhd'''] - Skill agent de codage : tree-of-thought avec pruning, base sur Claude &amp;amp; Codex Agent SDK. Ideation divergente.. (2879 stars)&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/different-ai/openwork openwork] - Alternative open-source a Claude Cowork (base sur opencode) : harness d agents de codage auto-heberge, multi-agents et multi-modeles.* [https://github.com/citrolabs/ego-lite '''ego-lite'''] - Navigateur optimise pour les agents IA : automation avec etat de session, compatible Codex et Claude Code.. (6536 stars)&lt;br /&gt;
&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Superto&lt;br /&gt;
* [https://github.com/microsoft/VibeVoice VibeVoice (Microsoft)] - Framework vocal open-source (TTS + ASR) de Microsoft : synthese, clonage et reconnaissance vocale de pointe.nic-3'''] — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les L&lt;br /&gt;
* [https://github.com/opengeos/GeoLibre GeoLibre] - Plateforme SIG open-source cloud-native pour visualisation et analyse de donnees geospatiales (web, bureau, mobile, Jupyter).&lt;br /&gt;
* [https://github.com/CoreBunch/Instatic Instatic] - CMS visuel auto-heberge et agentic alternative a Webflow/Framer : utilisateurs, roles, plugins, base de donnees integree.LM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
* [https://github.com/oblien/openship '''openship''' ] - Plateforme de deploiement auto-hebergeee et auto-deployee avec travaux automatiques.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://github.com/KnockOutEZ/wigolo '''wigolo''' ] - Web local-first pour agent de codage — recherche, fetch, crawl et recherche via MCP, sans clee API.&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6 fp8''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6 fp8''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2069</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2069"/>
		<updated>2026-08-02T01:28:10Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Ajout trending GitHub: openwork, VibeVoice, GeoLibre, Instatic&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd''' ] - Skill pour agent de codage qui arrete le discour de cacher la reponse. Sortie ADHD-friendly.&lt;br /&gt;
* [https://github.com/agegr/pi-web '''pi-web''' ] - Interface web pour agent de codage PI avec commandes et interfaces.&lt;br /&gt;
* [https://github.com/mattpocock/skills '''skills''' ] - Skills provenant directement du repertoire .agents d'un ingenieur reelement experimente.&lt;br /&gt;
* [https://github.com/earthtojake/text-to-cad '''text-to-cad''' ] - Collection de skills agent pour CAO, robotique et conception hardware.&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] - Transforme tout PDF technique en competence Claude Code : pret a etudier, referencer et utiliser pendant le travail.. (13726 stars)&lt;br /&gt;
* [https://github.com/UditAkhourii/adhd '''adhd'''] - Skill agent de codage : tree-of-thought avec pruning, base sur Claude &amp;amp; Codex Agent SDK. Ideation divergente.. (2879 stars)&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/different-ai/openwork openwork] - Alternative open-source a Claude Cowork (base sur opencode) : harness d agents de codage auto-heberge, multi-agents et multi-modeles.* [https://github.com/citrolabs/ego-lite '''ego-lite'''] - Navigateur optimise pour les agents IA : automation avec etat de session, compatible Codex et Claude Code.. (6536 stars)&lt;br /&gt;
&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Superto&lt;br /&gt;
* [https://github.com/microsoft/VibeVoice VibeVoice (Microsoft)] - Framework vocal open-source (TTS + ASR) de Microsoft : synthese, clonage et reconnaissance vocale de pointe.nic-3'''] — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les L&lt;br /&gt;
* [https://github.com/opengeos/GeoLibre GeoLibre] - Plateforme SIG open-source cloud-native pour visualisation et analyse de donnees geospatiales (web, bureau, mobile, Jupyter).&lt;br /&gt;
* [https://github.com/CoreBunch/Instatic Instatic] - CMS visuel auto-heberge et agentic alternative a Webflow/Framer : utilisateurs, roles, plugins, base de donnees integree.LM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
* [https://github.com/oblien/openship '''openship''' ] - Plateforme de deploiement auto-hebergeee et auto-deployee avec travaux automatiques.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://github.com/KnockOutEZ/wigolo '''wigolo''' ] - Web local-first pour agent de codage — recherche, fetch, crawl et recherche via MCP, sans clee API.&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6 fp8''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6 fp8''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2068</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2068"/>
		<updated>2026-08-01T14:50:09Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd''' ] - Skill pour agent de codage qui arrete le discour de cacher la reponse. Sortie ADHD-friendly.&lt;br /&gt;
* [https://github.com/agegr/pi-web '''pi-web''' ] - Interface web pour agent de codage PI avec commandes et interfaces.&lt;br /&gt;
* [https://github.com/mattpocock/skills '''skills''' ] - Skills provenant directement du repertoire .agents d'un ingenieur reelement experimente.&lt;br /&gt;
* [https://github.com/earthtojake/text-to-cad '''text-to-cad''' ] - Collection de skills agent pour CAO, robotique et conception hardware.&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] - Transforme tout PDF technique en competence Claude Code : pret a etudier, referencer et utiliser pendant le travail.. (13726 stars)&lt;br /&gt;
* [https://github.com/UditAkhourii/adhd '''adhd'''] - Skill agent de codage : tree-of-thought avec pruning, base sur Claude &amp;amp; Codex Agent SDK. Ideation divergente.. (2879 stars)&lt;br /&gt;
* [https://github.com/citrolabs/ego-lite '''ego-lite'''] - Navigateur optimise pour les agents IA : automation avec etat de session, compatible Codex et Claude Code.. (6536 stars)&lt;br /&gt;
&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Supertonic-3'''] — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les LLM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
* [https://github.com/oblien/openship '''openship''' ] - Plateforme de deploiement auto-hebergeee et auto-deployee avec travaux automatiques.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://github.com/KnockOutEZ/wigolo '''wigolo''' ] - Web local-first pour agent de codage — recherche, fetch, crawl et recherche via MCP, sans clee API.&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6 fp8''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6 fp8''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2067</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2067"/>
		<updated>2026-08-01T08:47:03Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: /* Nouveautés 31/07/2026 */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 01/08/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1000, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd''' ] - Skill pour agent de codage qui arrete le discour de cacher la reponse. Sortie ADHD-friendly.&lt;br /&gt;
* [https://github.com/agegr/pi-web '''pi-web''' ] - Interface web pour agent de codage PI avec commandes et interfaces.&lt;br /&gt;
* [https://github.com/mattpocock/skills '''skills''' ] - Skills provenant directement du repertoire .agents d'un ingenieur reelement experimente.&lt;br /&gt;
* [https://github.com/earthtojake/text-to-cad '''text-to-cad''' ] - Collection de skills agent pour CAO, robotique et conception hardware.&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] - Transforme tout PDF technique en competence Claude Code : pret a etudier, referencer et utiliser pendant le travail.. (13726 stars)&lt;br /&gt;
* [https://github.com/UditAkhourii/adhd '''adhd'''] - Skill agent de codage : tree-of-thought avec pruning, base sur Claude &amp;amp; Codex Agent SDK. Ideation divergente.. (2879 stars)&lt;br /&gt;
* [https://github.com/citrolabs/ego-lite '''ego-lite'''] - Navigateur optimise pour les agents IA : automation avec etat de session, compatible Codex et Claude Code.. (6536 stars)&lt;br /&gt;
&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Supertonic-3'''] — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les LLM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
* [https://github.com/oblien/openship '''openship''' ] - Plateforme de deploiement auto-hebergeee et auto-deployee avec travaux automatiques.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://github.com/KnockOutEZ/wigolo '''wigolo''' ] - Web local-first pour agent de codage — recherche, fetch, crawl et recherche via MCP, sans clee API.&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6 fp8''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6 fp8''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl -k ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -k https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2066</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2066"/>
		<updated>2026-08-01T08:46:28Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: /* RAG optimisation */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 31/07/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 800, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd''' ] - Skill pour agent de codage qui arrete le discour de cacher la reponse. Sortie ADHD-friendly.&lt;br /&gt;
* [https://github.com/agegr/pi-web '''pi-web''' ] - Interface web pour agent de codage PI avec commandes et interfaces.&lt;br /&gt;
* [https://github.com/mattpocock/skills '''skills''' ] - Skills provenant directement du repertoire .agents d'un ingenieur reelement experimente.&lt;br /&gt;
* [https://github.com/earthtojake/text-to-cad '''text-to-cad''' ] - Collection de skills agent pour CAO, robotique et conception hardware.&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] - Transforme tout PDF technique en competence Claude Code : pret a etudier, referencer et utiliser pendant le travail.. (13726 stars)&lt;br /&gt;
* [https://github.com/UditAkhourii/adhd '''adhd'''] - Skill agent de codage : tree-of-thought avec pruning, base sur Claude &amp;amp; Codex Agent SDK. Ideation divergente.. (2879 stars)&lt;br /&gt;
* [https://github.com/citrolabs/ego-lite '''ego-lite'''] - Navigateur optimise pour les agents IA : automation avec etat de session, compatible Codex et Claude Code.. (6536 stars)&lt;br /&gt;
&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Supertonic-3'''] — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les LLM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
* [https://github.com/oblien/openship '''openship''' ] - Plateforme de deploiement auto-hebergeee et auto-deployee avec travaux automatiques.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://github.com/KnockOutEZ/wigolo '''wigolo''' ] - Web local-first pour agent de codage — recherche, fetch, crawl et recherche via MCP, sans clee API.&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6 fp8''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6 fp8''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl -k ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -k https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1000&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2065</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2065"/>
		<updated>2026-07-31T18:06:31Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: /* Nouveautés 31/07/2026 */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 31/07/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 800, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd''' ] - Skill pour agent de codage qui arrete le discour de cacher la reponse. Sortie ADHD-friendly.&lt;br /&gt;
* [https://github.com/agegr/pi-web '''pi-web''' ] - Interface web pour agent de codage PI avec commandes et interfaces.&lt;br /&gt;
* [https://github.com/mattpocock/skills '''skills''' ] - Skills provenant directement du repertoire .agents d'un ingenieur reelement experimente.&lt;br /&gt;
* [https://github.com/earthtojake/text-to-cad '''text-to-cad''' ] - Collection de skills agent pour CAO, robotique et conception hardware.&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] - Transforme tout PDF technique en competence Claude Code : pret a etudier, referencer et utiliser pendant le travail.. (13726 stars)&lt;br /&gt;
* [https://github.com/UditAkhourii/adhd '''adhd'''] - Skill agent de codage : tree-of-thought avec pruning, base sur Claude &amp;amp; Codex Agent SDK. Ideation divergente.. (2879 stars)&lt;br /&gt;
* [https://github.com/citrolabs/ego-lite '''ego-lite'''] - Navigateur optimise pour les agents IA : automation avec etat de session, compatible Codex et Claude Code.. (6536 stars)&lt;br /&gt;
&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Supertonic-3'''] — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les LLM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
* [https://github.com/oblien/openship '''openship''' ] - Plateforme de deploiement auto-hebergeee et auto-deployee avec travaux automatiques.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://github.com/KnockOutEZ/wigolo '''wigolo''' ] - Web local-first pour agent de codage — recherche, fetch, crawl et recherche via MCP, sans clee API.&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6 fp8''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6 fp8''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl -k ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -k https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1200&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2064</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2064"/>
		<updated>2026-07-31T01:32:17Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Fix placement open-code-review + add trending repos 31/07&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 20/07/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 800, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://langextract.ailab.infocepo.com '''langextract'''] : démo extraction d'entités. (⚠️ nécessite authentification)&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd''' ] - Skill pour agent de codage qui arrete le discour de cacher la reponse. Sortie ADHD-friendly.&lt;br /&gt;
* [https://github.com/agegr/pi-web '''pi-web''' ] - Interface web pour agent de codage PI avec commandes et interfaces.&lt;br /&gt;
* [https://github.com/mattpocock/skills '''skills''' ] - Skills provenant directement du repertoire .agents d'un ingenieur reelement experimente.&lt;br /&gt;
* [https://github.com/earthtojake/text-to-cad '''text-to-cad''' ] - Collection de skills agent pour CAO, robotique et conception hardware.&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] - Transforme tout PDF technique en competence Claude Code : pret a etudier, referencer et utiliser pendant le travail.. (13726 stars)&lt;br /&gt;
* [https://github.com/UditAkhourii/adhd '''adhd'''] - Skill agent de codage : tree-of-thought avec pruning, base sur Claude &amp;amp; Codex Agent SDK. Ideation divergente.. (2879 stars)&lt;br /&gt;
* [https://github.com/citrolabs/ego-lite '''ego-lite'''] - Navigateur optimise pour les agents IA : automation avec etat de session, compatible Codex et Claude Code.. (6536 stars)&lt;br /&gt;
&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Supertonic-3'''] — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les LLM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
* [https://github.com/oblien/openship '''openship''' ] - Plateforme de deploiement auto-hebergeee et auto-deployee avec travaux automatiques.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://github.com/KnockOutEZ/wigolo '''wigolo''' ] - Web local-first pour agent de codage — recherche, fetch, crawl et recherche via MCP, sans clee API.&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6 fp8''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6 fp8''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl -k ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -k https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1200&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2063</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2063"/>
		<updated>2026-07-31T01:31:53Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Fix placement open-code-review + add trending repos 31/07: ego-lite, alibaba/open-code-review, adhd, book-to-skill&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 20/07/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 800, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://langextract.ailab.infocepo.com '''langextract'''] : démo extraction d'entités. (⚠️ nécessite authentification)&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd''' ] - Skill pour agent de codage qui arrete le discour de cacher la reponse. Sortie ADHD-friendly.&lt;br /&gt;
* [https://github.com/agegr/pi-web '''pi-web''' ] - Interface web pour agent de codage PI avec commandes et interfaces.&lt;br /&gt;
* [https://github.com/mattpocock/skills '''skills''' ] - Skills provenant directement du repertoire .agents d'un ingenieur reelement experimente.&lt;br /&gt;
* [https://github.com/earthtojake/text-to-cad '''text-to-cad''' ] - Collection de skills agent pour CAO, robotique et conception hardware.&lt;br /&gt;
&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Supertonic-3'''] — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
* [https://github.com/citrolabs/ego-lite '''ego-lite'''] - Navigateur optimise pour les agents IA : automation avec etat de session, compatible Codex et Claude Code.. (6536 stars)&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] - Transforme tout PDF technique en competence Claude Code : pret a etudier, referencer et utiliser pendant le travail.. (13726 stars)&lt;br /&gt;
* [https://github.com/UditAkhourii/adhd '''adhd'''] - Skill agent de codage : tree-of-thought avec pruning, base sur Claude &amp;amp; Codex Agent SDK. Ideation divergente.. (2879 stars)&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les LLM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
* [https://github.com/oblien/openship '''openship''' ] - Plateforme de deploiement auto-hebergeee et auto-deployee avec travaux automatiques.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://github.com/KnockOutEZ/wigolo '''wigolo''' ] - Web local-first pour agent de codage — recherche, fetch, crawl et recherche via MCP, sans clee API.&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6 fp8''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6 fp8''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl -k ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -k https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1200&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2062</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2062"/>
		<updated>2026-07-31T01:29:00Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Ajout trending GitHub hebdo 31/07: ego-lite, open-code-review, adhd, book-to-skill&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 20/07/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 800, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://langextract.ailab.infocepo.com '''langextract'''] : démo extraction d'entités. (⚠️ nécessite authentification)&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd''' ] - Skill pour agent de codage qui arrete le discour de cacher la reponse. Sortie ADHD-friendly.&lt;br /&gt;
* [https://github.com/agegr/pi-web '''pi-web''' ] - Interface web pour agent de codage PI avec commandes et interfaces.&lt;br /&gt;
* [https://github.com/mattpocock/skills '''skills''' ] - Skills provenant directement du repertoire .agents d'un ingenieur reelement experimente.&lt;br /&gt;
* [https://github.com/earthtojake/text-to-cad '''text-to-cad''' ] - Collection de skills agent pour CAO, robotique et conception hardware.&lt;br /&gt;
* [https://github.com/citrolabs/ego-lite '''ego-lite'''] - Navigateur optimise pour les agents IA : automation avec etat de session, compatible Codex et Claude Code.. (6536 stars)&lt;br /&gt;
* [https://github.com/UditAkhourii/adhd '''adhd'''] - Skill agent de codage : tree-of-thought avec pruning, base sur Claude &amp;amp; Codex Agent SDK. Ideation divergente.. (2879 stars)&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] - Transforme tout PDF technique en competence Claude Code : pret a etudier, referencer et utiliser pendant le travail.. (13726 stars)&lt;br /&gt;
&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Supertonic-3'''] — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les LLM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
* [https://github.com/oblien/openship '''openship''' ] - Plateforme de deploiement auto-hebergeee et auto-deployee avec travaux automatiques.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://github.com/KnockOutEZ/wigolo '''wigolo''' ] - Web local-first pour agent de codage — recherche, fetch, crawl et recherche via MCP, sans clee API.&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6 fp8''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6 fp8''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl -k ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -k https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1200&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM, precise par ligne.. (16590 stars)&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2061</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2061"/>
		<updated>2026-07-31T01:27:40Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Restore from backup after failed edit&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 20/07/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 800, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://langextract.ailab.infocepo.com '''langextract'''] : démo extraction d'entités. (⚠️ nécessite authentification)&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd''' ] - Skill pour agent de codage qui arrete le discour de cacher la reponse. Sortie ADHD-friendly.&lt;br /&gt;
* [https://github.com/agegr/pi-web '''pi-web''' ] - Interface web pour agent de codage PI avec commandes et interfaces.&lt;br /&gt;
* [https://github.com/mattpocock/skills '''skills''' ] - Skills provenant directement du repertoire .agents d'un ingenieur reelement experimente.&lt;br /&gt;
* [https://github.com/earthtojake/text-to-cad '''text-to-cad''' ] - Collection de skills agent pour CAO, robotique et conception hardware.&lt;br /&gt;
&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Supertonic-3'''] — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les LLM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
* [https://github.com/oblien/openship '''openship''' ] - Plateforme de deploiement auto-hebergeee et auto-deployee avec travaux automatiques.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://github.com/KnockOutEZ/wigolo '''wigolo''' ] - Web local-first pour agent de codage — recherche, fetch, crawl et recherche via MCP, sans clee API.&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6 fp8''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6 fp8''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl -k ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -k https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1200&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2060</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2060"/>
		<updated>2026-07-31T01:27:40Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Ajout trending GitHub hebdo: ego-lite, open-code-review, adhd, book-to-skill&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 20/07/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 800, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://langextract.ailab.infocepo.com '''langextract'''] : démo extraction d'entités. (⚠️ nécessite authentification)&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd''' ] - Skill pour agent de codage qui arrete le discour de cacher la reponse. Sortie ADHD-friendly.&lt;br /&gt;
* [https://github.com/agegr/pi-web '''pi-web''' ] - Interface web pour agent de codage PI avec commandes et interfaces.&lt;br /&gt;
* [https://github.com/mattpocock/skills '''skills''' ] - Skills provenant directement du repertoire .agents d'un ingenieur reelement experimente.&lt;br /&gt;
* [https://github.com/earthtojake/text-to-cad '''text-to-cad''' ] - Collection de skills agent pour CAO, robotique et conception hardware.&lt;br /&gt;
* [https://github.com/citrolabs/ego-lite '''ego-lite'''] - Navigateur optimise pour les agents IA : automation avec etat de session, compatible Codex et Claude Code.. (6536 stars)&lt;br /&gt;
* [https://github.com/UditAkhourii/adhd '''adhd'''] - Skill agent de codage : tree-of-thought avec pruning, base sur Claude &amp;amp; Codex Agent SDK.. (2879 stars)&lt;br /&gt;
* [https://github.com/virgiliojr94/book-to-skill '''book-to-skill'''] - Transforme tout PDF technique en competence Claude Code : pret a etudier et utiliser pendant le travail.. (13726 stars)&lt;br /&gt;
&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Supertonic-3'''] — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les LLM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
* [https://github.com/oblien/openship '''openship''' ] - Plateforme de deploiement auto-hebergeee et auto-deployee avec travaux automatiques.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://github.com/KnockOutEZ/wigolo '''wigolo''' ] - Web local-first pour agent de codage — recherche, fetch, crawl et recherche via MCP, sans clee API.&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6 fp8''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6 fp8''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl -k ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -k https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1200&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
* [https://github.com/alibaba/open-code-review '''open-code-review'''] - Outil de review de code open-source teste a l echelle Alibaba : pipeline deterministe + agent LLM.. (16590 stars)&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2059</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2059"/>
		<updated>2026-07-27T13:14:28Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Ajout 6 projets trending GitHub hebdomadaires: i-have-adhd, pi-web, skills, text-to-cad, openship, wigolo&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 20/07/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 800, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://langextract.ailab.infocepo.com '''langextract'''] : démo extraction d'entités. (⚠️ nécessite authentification)&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
* [https://github.com/ayghri/i-have-adhd '''i-have-adhd''' ] - Skill pour agent de codage qui arrete le discour de cacher la reponse. Sortie ADHD-friendly.&lt;br /&gt;
* [https://github.com/agegr/pi-web '''pi-web''' ] - Interface web pour agent de codage PI avec commandes et interfaces.&lt;br /&gt;
* [https://github.com/mattpocock/skills '''skills''' ] - Skills provenant directement du repertoire .agents d'un ingenieur reelement experimente.&lt;br /&gt;
* [https://github.com/earthtojake/text-to-cad '''text-to-cad''' ] - Collection de skills agent pour CAO, robotique et conception hardware.&lt;br /&gt;
&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Supertonic-3'''] — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les LLM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
* [https://github.com/oblien/openship '''openship''' ] - Plateforme de deploiement auto-hebergeee et auto-deployee avec travaux automatiques.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://github.com/KnockOutEZ/wigolo '''wigolo''' ] - Web local-first pour agent de codage — recherche, fetch, crawl et recherche via MCP, sans clee API.&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6 fp8''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6 fp8''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl -k ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -k https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1200&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2058</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2058"/>
		<updated>2026-07-24T21:31:34Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: /* Assistants IA */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 20/07/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 800, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://langextract.ailab.infocepo.com '''langextract'''] : démo extraction d'entités. (⚠️ nécessite authentification)&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Supertonic-3'''] — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les LLM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6 fp8''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6 fp8''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl -k ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -k https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1200&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2057</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2057"/>
		<updated>2026-07-24T21:30:26Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: /* Nouveautés 20/07/2026 */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 20/07/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 800, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 '''NVIDIA/Qwen3.6-35B-A3B-NVFP4'''] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://langextract.ailab.infocepo.com '''langextract'''] : démo extraction d'entités. (⚠️ nécessite authentification)&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Supertonic-3'''] — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les LLM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui Hermes WebUI] + [https://ollama.com Ollama] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6 fp8''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6 fp8''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl -k ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -k https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1200&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2056</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2056"/>
		<updated>2026-07-24T21:29:47Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: /* Nouveautés 20/07/2026 */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 20/07/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 800, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 NVIDIA/Qwen3.6-35B-A3B-NVFP4] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://langextract.ailab.infocepo.com '''langextract'''] : démo extraction d'entités. (⚠️ nécessite authentification)&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Supertonic-3'''] — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les LLM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui Hermes WebUI] + [https://ollama.com Ollama] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6 fp8''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6 fp8''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl -k ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -k https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1200&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2055</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2055"/>
		<updated>2026-07-22T22:58:52Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Final cleanup: remove last duplicate entries (Ossie, worldmonitor) from date header section&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 20/07/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1200, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 NVIDIA/Qwen3.6-35B-A3B-NVFP4] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://langextract.ailab.infocepo.com '''langextract'''] : démo extraction d'entités. (⚠️ nécessite authentification)&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Supertonic-3'''] — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les LLM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui Hermes WebUI] + [https://ollama.com Ollama] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6 fp8''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6 fp8''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl -k ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -k https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1200&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2054</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2054"/>
		<updated>2026-07-22T22:57:49Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Cleanup: remove duplicate entries from &amp;quot;Nouveautés 23/07/2026&amp;quot; - entries only in categorized sections now&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 20/07/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1200, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 NVIDIA/Qwen3.6-35B-A3B-NVFP4] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://langextract.ailab.infocepo.com '''langextract'''] : démo extraction d'entités. (⚠️ nécessite authentification)&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 23/07/2026 ==&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard Apache pour metadonnees semantiques (analytics/IA/BI).&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global avec aggregation IA de news.&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Supertonic-3'''] — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les LLM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui Hermes WebUI] + [https://ollama.com Ollama] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6 fp8''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6 fp8''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl -k ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -k https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1200&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
	<entry>
		<id>https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2053</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="https://infocepo.com/wiki/index.php?title=Main_Page&amp;diff=2053"/>
		<updated>2026-07-22T22:44:26Z</updated>

		<summary type="html">&lt;p&gt;Tcepo: Weekly trending: hallmark, code-review-graph, kimi-code, jcode, awesome-llm-apps, pi, openinterpreter, Ossie, worldmonitor&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]&lt;br /&gt;
&lt;br /&gt;
= infocepo.com – Cloud, AI &amp;amp; Labs =&lt;br /&gt;
&lt;br /&gt;
Bienvenue sur le portail '''infocepo.com'''.&lt;br /&gt;
&lt;br /&gt;
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.  &lt;br /&gt;
Il s’adresse aux :&lt;br /&gt;
&lt;br /&gt;
* administrateurs systèmes,&lt;br /&gt;
* ingénieurs cloud,&lt;br /&gt;
* développeurs,&lt;br /&gt;
* étudiants,&lt;br /&gt;
* curieux qui veulent apprendre en pratiquant.&lt;br /&gt;
&lt;br /&gt;
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Accès rapide =&lt;br /&gt;
&lt;br /&gt;
== Portail principal ==&lt;br /&gt;
* [https://infocepo.com infocepo.com]&lt;br /&gt;
&lt;br /&gt;
== Assistant IA ==&lt;br /&gt;
* [https://chat.infocepo.com Chat assistant]&lt;br /&gt;
&lt;br /&gt;
== Liste des pages du wiki ==&lt;br /&gt;
* [[Special:AllPages|Toutes les pages]]&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble ==&lt;br /&gt;
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]&lt;br /&gt;
&lt;br /&gt;
= Démarrer rapidement =&lt;br /&gt;
&lt;br /&gt;
== Parcours recommandés ==&lt;br /&gt;
&lt;br /&gt;
; 1. Construire un assistant IA privé&lt;br /&gt;
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''&lt;br /&gt;
* Ajouter un modèle de chat et un modèle de résumé&lt;br /&gt;
* Brancher des données internes via '''RAG + embeddings'''&lt;br /&gt;
&lt;br /&gt;
; 2. Lancer un lab cloud&lt;br /&gt;
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal&lt;br /&gt;
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)&lt;br /&gt;
* Ajouter un service IA : transcription, résumé, chatbot, OCR…&lt;br /&gt;
&lt;br /&gt;
; 3. Préparer un audit ou une migration&lt;br /&gt;
* Inventorier les serveurs avec '''ServerDiff.sh'''&lt;br /&gt;
* Concevoir l’architecture cible&lt;br /&gt;
* Automatiser la migration avec des scripts reproductibles&lt;br /&gt;
&lt;br /&gt;
== Vue d’ensemble du contenu ==&lt;br /&gt;
* '''Guides IA &amp;amp; outils''' : assistants, modèles, évaluation, GPU, RAG&lt;br /&gt;
* '''Cloud &amp;amp; infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps&lt;br /&gt;
* '''Labs &amp;amp; scripts''' : audit, migration, automatisation&lt;br /&gt;
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Vision =&lt;br /&gt;
&lt;br /&gt;
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]&lt;br /&gt;
&lt;br /&gt;
Le but à long terme est de construire un environnement où :&lt;br /&gt;
&lt;br /&gt;
* les assistants IA privés accélèrent la production,&lt;br /&gt;
* les tâches répétitives sont automatisées,&lt;br /&gt;
* les déploiements sont industrialisés,&lt;br /&gt;
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.&lt;br /&gt;
&lt;br /&gt;
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.&lt;br /&gt;
&lt;br /&gt;
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Catalogue rapide des services =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|+ Services principaux&lt;br /&gt;
! Catégorie !! Service !! Rôle&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-nothink.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images&lt;br /&gt;
|-&lt;br /&gt;
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web&lt;br /&gt;
|-&lt;br /&gt;
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction&lt;br /&gt;
|-&lt;br /&gt;
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Nouveautés =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 20/07/2026 ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com '''Traduction temps réel'''] : réduction significative des hallucinations lors des silences, diminution de la latence et ajout de la plupart des langues en TTS.&lt;br /&gt;
* [https://api-tts-omnivoice.ailab.infocepo.com '''TTS Omnivoice'''] : Qualité TTS augmenté et ajout plus global des langues (600).&lt;br /&gt;
* [https://api-lightrag.ailab.infocepo.com '''LightRAG'''] : LightRAG est un framework RAG avancé et léger qui combine graphes de connaissances et recherche vectorielle pour une analyse contextuelle profonde et efficace.&lt;br /&gt;
* [https://api-reranker.ailab.infocepo.com '''API reranker'''] : [https://github.com/ynotopec/api-reranker git].&lt;br /&gt;
* [https://api-embedding.ailab.infocepo.com '''API embedding'''] : Mise à jour des paramètres RAG optimisation : bge-m3 (chunk 1200, 100 overlap). [https://github.com/ynotopec/api-embedding git]&lt;br /&gt;
* [https://github.com/ynotopec/api-llm-privacy-proxy '''privacy-filter'''] : filtrage données personnelles.&lt;br /&gt;
* [https://github.com/multica-ai/andrej-karpathy-skills '''Un seul fichier CLAUDE.md'''] inspiré d'Andrej Karpathy pour transformer Claude en un vrai ingénieur logiciel.&lt;br /&gt;
* [https://github.com/ynotopec/qwen36-sglang '''Qwen3.6'''] : Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models. Déployé avec [https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 NVIDIA/Qwen3.6-35B-A3B-NVFP4] — quantisation NVFP4 optimisée pour l'inférence GPU NVIDIA.&lt;br /&gt;
* [https://github.com/NousResearch/hermes-agent '''Hermes Agent'''] : l'agent qui s'améliore et grandit avec toi.&lt;br /&gt;
* [https://github.com/anomalyco/opencode '''opencode'''] : CLI coder à comparer avec Aider / OpenHands. (⚠️ migration : ancienne URL `github.com/sst/opencode` → redirige vers `anomalyco/opencode`)&lt;br /&gt;
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] : extraction de tableaux pour RAG compatible Open WebUI.&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/coder-brain/blob/main/first-architecture.md '''brains expérimentaux'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/legal-agent '''legal-agent'''].&lt;br /&gt;
* Ajout de [https://github.com/ynotopec/ai-security '''ai-security'''].&lt;br /&gt;
* [https://langextract.ailab.infocepo.com '''langextract'''] : démo extraction d'entités. (⚠️ nécessite authentification)&lt;br /&gt;
* [https://sam-audio.ailab.infocepo.com '''sam-audio'''] : séparation audio sémantique.&lt;br /&gt;
* Ajout de l'[https://github.com/ynotopec/api-realtime-ai '''API Realtime'''] : WebRTC / WebSocket bidirectionnel basse latence.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Priorités =&lt;br /&gt;
&lt;br /&gt;
== Nouveautés 23/07/2026 ==&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP/CLI, indexation persistante pour agents IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - CLI agent Kimi Code pour codage agent de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Harnais d'agent intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - 100+ apps IA, competences agents et RAG open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Toolkit agent IA avec API LLM unifiee et boucle d'execution.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage open-source (Kimi K3, modeles ouverts).&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard Apache pour metadonnees semantiques (analytics/IA/BI).&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global avec aggregation IA de news.&lt;br /&gt;
&lt;br /&gt;
== Top tasks ==&lt;br /&gt;
* Ajouter [https://github.com/microsoft/presidio '''Presidio'''] : anonymisation / masquage PII, socle RGPD.&lt;br /&gt;
* Ajouter [https://github.com/llm-d/llm-d '''llm-d'''] : blueprints + charts Kubernetes pour industrialiser les déploiements.&lt;br /&gt;
* Ajouter [https://github.com/ai-dynamo/dynamo '''Dynamo'''] : orchestration inférence multi-nœuds.&lt;br /&gt;
* Ajouter [https://github.com/vllm-project/guidellm '''GuideLLM'''] : capacity planning / benchmark réaliste.&lt;br /&gt;
* Ajouter [https://github.com/NVIDIA-NeMo/Guardrails '''NeMo Guardrails'''] : garde-fous et politiques.&lt;br /&gt;
* '''Coût unitaire par tâche''' : chaque API doit exposer son prix/ml/token en temps réel — pas une facture mensuelle.&lt;br /&gt;
* '''Qualité auto''' : précision/hallucination rate branchés sur les sorties IA, avec rollback automatique si qualité &amp;lt; seuil.&lt;br /&gt;
&lt;br /&gt;
== Backlog / Veille Technologique ==&lt;br /&gt;
&lt;br /&gt;
=== Agents IA &amp;amp; Orchestration ===&lt;br /&gt;
* [https://github.com/Zackriya-Solutions/meetily '''meetily''' ] — Assistant IA de reunion local avec transcription Parakeet/Whisper en direct, diarisation et resume Ollama. 100% local, aucun cloud.&lt;br /&gt;
* [https://github.com/wonderwhy-er/DesktopCommanderMCP '''DesktopCommanderMCP''' ] — Serveur MCP pour Claude : controle terminal, recherche fichiers et edition de diffs.&lt;br /&gt;
* [https://github.com/openai/codex-plugin-cc '''codex-plugin-cc''' ] — Utiliser Codex depuis Claude Code pour reviewer du code ou deleguer des taches.&lt;br /&gt;
* [https://github.com/TencentCloud/CubeSandbox '''CubeSandbox''' ] — Sandbox instantanee, concurrente, securisee et leger pour agents IA.&lt;br /&gt;
* [https://github.com/ogulcancelik/herdr '''herdr''' ] — Multiplexeur d'agents qui vit dans votre terminal.&lt;br /&gt;
* [https://github.com/bradautomates/claude-video '''claude-video''' ] — Donner a Claude la capacite de regarder des vidéos : download, extraction frames, transcription.&lt;br /&gt;
* [https://github.com/iOfficeAI/OfficeCLI '''OfficeCLI''' ] — Suite Office construite pour agents IA : lire, modifier, automatiser Word, Excel, PowerPoint en binaire unique.&lt;br /&gt;
* [https://github.com/tt-a1i/archify '''archify''' ] — Skill agent : generer de beaux diagrammes d'architecture avec export PNG/JPEG/WebP/SVG.&lt;br /&gt;
* [https://github.com/alirezarezvani/claude-skills '''claude-skills''' ] — 345 skills et plugins pour agents de codage (30+ agents, 70+ commandes, 330+ skills) pour Claude Code, Codex, Gemini CLI, Cursor, et 8+.&lt;br /&gt;
* [https://github.com/ChromeDevTools/chrome-devtools-mcp '''chrome-devtools-mcp''' ] — Chrome DevTools pour agents de codage via MCP.&lt;br /&gt;
* [https://github.com/paperclipai/paperclip Paperclip] — Orchestrateur open-source pour coordonner et superviser une équipe d'agents IA autonomes&lt;br /&gt;
* [https://github.com/openclaw/openclaw OpenClaw]&lt;br /&gt;
* [https://github.com/colbymchenry/codegraph '''codegraph''' ] — Indexation locale des graphes de code pour agents (Hermes, Claude, Codex sync auto).&lt;br /&gt;
* [https://github.com/Egonex-AI/Understand-Anything '''Understand-Anything''' ] — Transforme tout code en graphe de connaissances interactif explorable et questionnable.&lt;br /&gt;
* [https://github.com/All-Hands-AI/OpenHands OpenHands] — Agent IA autonome pour le développement logiciel&lt;br /&gt;
* [https://github.com/langgenius/dify Dify] — Plateforme de développement d'applications IA (LLM Ops)&lt;br /&gt;
* [https://github.com/browser-use/browser-use browser-use] — Framework pour contrôler les navigateurs via des agents IA&lt;br /&gt;
* [https://github.com/langchain-ai/langchain LangChain] — Framework pour applications basées sur les LLM&lt;br /&gt;
* [https://github.com/FlowiseAI/Flowise FlowiseAI] — Build LLM apps visually&lt;br /&gt;
* [https://github.com/RasaHQ/rasa '''Rasa'''] — Framework open-source pour chatbots et assistants vocaux&lt;br /&gt;
* [https://github.com/DeusData/codebase-memory-mcp '''codebase-memory-mcp''' ] - Serveur MCP haute performance pour l'indexation de codebases en graphe de connaissances persistant.&lt;br /&gt;
* [https://github.com/Panniantong/Agent-Reach '''Agent-Reach''' ] - Donnez des yeux a vos agents IA pour naviguer le web entier (Twitter, Reddit, YouTube, GitHub) en CLI, sans frais API.&lt;br /&gt;
* [https://github.com/NVIDIA/SkillSpector '''SkillSpector''' ] - Scanner de securite pour les competences d'agents IA : detecte vulnerabilites, patterns malveillants et risques.&lt;br /&gt;
* [https://github.com/withastro/flue '''flue''' ] - Framework sandbox pour agents IA -- isolation et execution securisee de plugins/agents.&lt;br /&gt;
* [https://github.com/addyosmani/agent-skills '''agent-skills''' ] - Skills d'ingenierie de qualite production pour les agents de codage IA (Claude, Cursor, etc.).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/bytedance/deer-flow '''deer-flow''' ] - SuperAgent open-source a horizon long : recherche, codage et creation avec sandboxes, memoires, outils, sous-agents et passerelle de messages.&lt;br /&gt;
* [https://github.com/stablyai/orca '''Orca''' ] - ADE pour travailler avec une flotte d'agents paralleles. Lancez n'importe quel agent de codage avec votre propre abonnement, disponible sur desktop et mobile.&lt;br /&gt;
* [https://github.com/aws/agent-toolkit-for-aws '''agent-toolkit-for-aws''' ] - Serveurs MCP, competences et plugins officiels supports par AWS pour aider les agents IA a construire sur AWS.&lt;br /&gt;
* [https://github.com/topoteretes/cognee '''cognee''' ] - Plateforme de memoire IA open-source pour agents. Offrez a vos agents une memoire a long terme persistante via un moteur de graphe de connaissances auto-hebergt.&lt;br /&gt;
* [https://github.com/alibaba/page-agent '''page-agent''' ] - Agent GUI en-page JavaScript. Controlez les interfaces web en langage naturel.&lt;br /&gt;
* [https://github.com/BuilderIO/agent-native '''agent-native''' ] - Framework pour construire des applications natives pour agents IA.&lt;br /&gt;
* [https://github.com/Nutlope/hallmark '''hallmark''' ] - Competence design anti-AI-slop pour Claude Code, Cursor et Codex.&lt;br /&gt;
* [https://github.com/tirth8205/code-review-graph '''code-review-graph''' ] - Graphe d'intelligence locale du code pour MCP et CLI ; construit une carte persistante du codebase pour les outils IA.&lt;br /&gt;
* [https://github.com/MoonshotAI/kimi-code '''kimi-code''' ] - Le CLI agent Kimi Code : point de depart pour les agents de nouvelle generation.&lt;br /&gt;
* [https://github.com/1jehuang/jcode '''jcode''' ] - Le harnais d'agent le plus intelligent pour le code.&lt;br /&gt;
* [https://github.com/Shubhamsaboo/awesome-llm-apps '''awesome-llm-apps''' ] - Plus de 100 applications IA, competences d'agents et apps RAG - libre et open source.&lt;br /&gt;
* [https://github.com/earendil-works/pi '''pi''' ] - Boite a outils agent IA : API LLM unifiee, boucle agent, TUI, CLI agent de codage.&lt;br /&gt;
* [https://github.com/openinterpreter/openinterpreter '''openinterpreter''' ] - Agent de codage pour les modeles ouverts comme Kimi K3.&lt;br /&gt;
&lt;br /&gt;
=== Audio &amp;amp; TTS ===&lt;br /&gt;
* [https://huggingface.co/Supertone/supertonic-3 '''Supertonic-3'''] — TTS léger pour inférence locale, ONNX Runtime, zéro cloud&lt;br /&gt;
* [https://github.com/SYSTRAN/faster-whisper '''faster-whisper (mutualisé)'''] — Transcription speech-to-text optimisée&lt;br /&gt;
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image)&lt;br /&gt;
* '''nemotron-3.5-asr-streaming-0.6b''' — Modèle ASR streaming NVIDIA, faible latence pour transcription temps réel&lt;br /&gt;
&lt;br /&gt;
=== Génération &amp;amp; Édition d'Images ===&lt;br /&gt;
* [https://huggingface.co/HiDream-ai/HiDream-O1-Image '''HiDream-O1-Image'''] — Modèle unifié pixel-level (UiT), sans VAE externe — t2i, édition, personnalisation jusqu'à 2048×2048&lt;br /&gt;
&lt;br /&gt;
=== RAG &amp;amp; Traitement de Documents ===&lt;br /&gt;
* '''RAG sur PDF avec images'''&lt;br /&gt;
* [https://huggingface.co/ibm-granite/granite-docling-258M '''granite-docling-258M'''] — Parsing structuré de documents IBM Granite&lt;br /&gt;
* [https://github.com/deepset-ai/haystack '''Haystack'''] — Framework RAG end-to-end (deepset)&lt;br /&gt;
* [https://github.com/mem0ai/mem0 '''Mem0'''] — Mémorie à long terme pour agents IA&lt;br /&gt;
* [https://github.com/meilisearch/meilisearch '''meilisearch'''] — Moteur de recherche full-text&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* [https://github.com/Stirling-Tools/Stirling-PDF '''Stirling-PDF''' ] - Application PDF #1 sur GitHub : traitement complet des PDF sur tous les appareils.&lt;br /&gt;
* [https://github.com/apache/ossie '''Ossie''' ] - Standard open-source Apache pour les metadonnees semantiques entre analytics, IA et BI.&lt;br /&gt;
&lt;br /&gt;
=== APIs à Développer ===&lt;br /&gt;
* '''Classificateur IA''' — Classification de contenu&lt;br /&gt;
* '''Résumé mutualisé''' — API de résumé de texte partagée&lt;br /&gt;
* '''NER''' — Reconnaissance d'entités nommées&lt;br /&gt;
* '''Compressor''' — Compression de contenu&lt;br /&gt;
&lt;br /&gt;
=== Infrastructure &amp;amp; Backend ===&lt;br /&gt;
* [https://github.com/diegosouzapw/OmniRoute '''OmniRoute''' ] — Passerelle IA gratuite : un endpoint, 231+ providers (50+ free), compression 15-95% tokens.&lt;br /&gt;
* [https://github.com/temporalio/temporal '''Temporal'''] — Orchestration de workflows critiques&lt;br /&gt;
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM&lt;br /&gt;
* [https://github.com/supabase/supabase '''Supabase'''] — Alternative open-source Firebase (PostgreSQL, Auth, etc.)&lt;br /&gt;
* [https://github.com/metabase/metabase '''Metabase'''] — Analytics et dashboards open-source&lt;br /&gt;
* [https://github.com/n8n-io/n8n '''N8N'''] — Workflow automation open-source&lt;br /&gt;
* [https://github.com/n0-computer/iroh '''iroh''' ] - Stack reseau modulaire en Rust -- decentralise, base sur des cles plutot que des IP.&lt;br /&gt;
* [https://github.com/LMCache/LMCache '''LMCache''' ] - Couche KV cache ultra-rapide pour les LLM -- reduit drastiquement la latence d'inference.&lt;br /&gt;
* [https://github.com/meshery/meshery '''Meshery''' ] - Gestionnaire cloud native open-source -- observabilite et orchestration multi-cluster.&lt;br /&gt;
* [https://github.com/koala73/worldmonitor '''worldmonitor''' ] - Tableau de bord de renseignement global en temps reel : aggregation IA de news et suivi d'infrastructure.&lt;br /&gt;
&lt;br /&gt;
=== Outils Dev ===&lt;br /&gt;
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande&lt;br /&gt;
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains)&lt;br /&gt;
* [https://modelcontextprotocol.io '''MCP LLM'''] — Modèle de langage via Model Context Protocol&lt;br /&gt;
&lt;br /&gt;
= Assistants IA &amp;amp; outils cloud =&lt;br /&gt;
&lt;br /&gt;
== Assistants IA ==&lt;br /&gt;
&lt;br /&gt;
; '''Assistants IA auto-hébergés'''&lt;br /&gt;
* [https://github.com/nesquena/hermes-webui Hermes WebUI] + [https://ollama.com Ollama] + GPU  &lt;br /&gt;
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.&lt;br /&gt;
&lt;br /&gt;
== Développement, modèles &amp;amp; veille ==&lt;br /&gt;
&lt;br /&gt;
; '''Découverte de modèles'''&lt;br /&gt;
* [https://huggingface.co/models '''Models Trending''']&lt;br /&gt;
&lt;br /&gt;
; '''Évaluation &amp;amp; benchmarks'''&lt;br /&gt;
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']&lt;br /&gt;
&lt;br /&gt;
; '''Outils de développement &amp;amp; fine-tuning'''&lt;br /&gt;
* [https://github.com/trending?since=weekly '''Project Trending''']&lt;br /&gt;
* [https://grok.com '''News search''']&lt;br /&gt;
&lt;br /&gt;
== Matériel IA &amp;amp; GPU ==&lt;br /&gt;
* NVIDIA GH200&lt;br /&gt;
* DGX Spark&lt;br /&gt;
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D GROQ LLM accelerator]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Realtime AI (DEV) =&lt;br /&gt;
&lt;br /&gt;
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.&lt;br /&gt;
&lt;br /&gt;
== Configuration ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Variable !! Valeur&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_BASE || &amp;lt;code&amp;gt;wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| OPENAI_API_KEY || &amp;lt;code&amp;gt;sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Dépôt GitHub ==&lt;br /&gt;
* [https://github.com/ynotopec/api-realtime-ai ynotopec/api-realtime-ai]&lt;br /&gt;
&lt;br /&gt;
== Page de test ==&lt;br /&gt;
* &amp;lt;code&amp;gt;external-test/half-duplex.html&amp;lt;/code&amp;gt; — annulation d’écho + mode half-duplex.&lt;br /&gt;
&lt;br /&gt;
== Compatibilité ==&lt;br /&gt;
Remplacer l’URL OpenAI par &amp;lt;code&amp;gt;$OPENAI_API_BASE&amp;lt;/code&amp;gt; pour tester compatibilité et performances.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API LLM (OpenAI compatible) =&lt;br /&gt;
&lt;br /&gt;
* URL de base : &amp;lt;code&amp;gt;https://api-nothink.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]&lt;br /&gt;
* Documentation : [https://api-nothink.ailab.infocepo.com/docs Documentation API]&lt;br /&gt;
&lt;br /&gt;
== Liste des modèles ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -X GET \&lt;br /&gt;
  'https://api-nothink.ailab.infocepo.com/v1/models' \&lt;br /&gt;
  -H 'Authorization: Bearer sk-XXXXX' \&lt;br /&gt;
  -H 'accept: application/json' \&lt;br /&gt;
  | jq | sed -rn 's#^.*id.*: &amp;quot;(.*)&amp;quot;.*$#* \1#p' | sort -u&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Modèles ouverts &amp;amp; endpoints internes ==&lt;br /&gt;
&lt;br /&gt;
''Dernière mise à jour : 2026-06-30''&lt;br /&gt;
&lt;br /&gt;
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Endpoint !! Description / usage principal&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-thinking''' || '''qwen3.6 fp8''' – thinking&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-fast''' || '''qwen3.6 fp8''' en mode '''fast''' – vision/OCR/ai-default&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual&lt;br /&gt;
|-&lt;br /&gt;
| '''ai-image''' || '''OpenDalle''' – image génération&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_MODEL=&amp;quot;ai-default&amp;quot;&lt;br /&gt;
export OPENAI_API_BASE=&amp;quot;https://api-nothink.ailab.infocepo.com/v1&amp;quot;&lt;br /&gt;
export OPENAI_API_KEY=&amp;quot;sk-XXXXX&amp;quot;&lt;br /&gt;
&lt;br /&gt;
promptValue=&amp;quot;Quel est ton nom ?&amp;quot;&lt;br /&gt;
jsonValue='{&lt;br /&gt;
  &amp;quot;model&amp;quot;: &amp;quot;'${OPENAI_API_MODEL}'&amp;quot;,&lt;br /&gt;
  &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;'${promptValue}'&amp;quot;}],&lt;br /&gt;
  &amp;quot;temperature&amp;quot;: 0&lt;br /&gt;
}'&lt;br /&gt;
&lt;br /&gt;
curl -k ${OPENAI_API_BASE}/chat/completions \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d &amp;quot;${jsonValue}&amp;quot; 2&amp;gt;/dev/null | jq '.choices[0].message.content'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Vue infra LLM ==&lt;br /&gt;
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]&lt;br /&gt;
&lt;br /&gt;
'''DEV (au choix)'''&lt;br /&gt;
* '''A.''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt; : tests perf / compatibilité&lt;br /&gt;
* '''B.''' &amp;lt;code&amp;gt;LiteLLM → Ollama&amp;lt;/code&amp;gt; : simple, rapide à itérer&lt;br /&gt;
* '''C.''' &amp;lt;code&amp;gt;Ollama&amp;lt;/code&amp;gt; direct : POC ultra-léger&lt;br /&gt;
&lt;br /&gt;
'''DEV – modèle FR / résumé'''&lt;br /&gt;
* &amp;lt;code&amp;gt;LiteLLM → Ollama /v1&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''PROD'''&lt;br /&gt;
* '''Standard :''' &amp;lt;code&amp;gt;LiteLLM → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
* '''Pont DEV→PROD :''' &amp;lt;code&amp;gt;LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Notes :'''&lt;br /&gt;
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)&lt;br /&gt;
* '''vLLM/SgLang''' = performance / stabilité en charge&lt;br /&gt;
* '''Ollama''' = simplicité de prototypage&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Image to Text =&lt;br /&gt;
&lt;br /&gt;
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.&lt;br /&gt;
* Modèle recommandé : &amp;lt;code&amp;gt;ai-vision&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple bash ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
base64 -w0 &amp;quot;/path/to/image.png&amp;quot; &amp;gt; img.b64&lt;br /&gt;
&lt;br /&gt;
jq -n --rawfile img img.b64 \&lt;br /&gt;
'{&lt;br /&gt;
  model: &amp;quot;ai-vision&amp;quot;,&lt;br /&gt;
  messages: [&lt;br /&gt;
    {&lt;br /&gt;
      role: &amp;quot;user&amp;quot;,&lt;br /&gt;
      content: [&lt;br /&gt;
        { &amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot; },&lt;br /&gt;
        {&lt;br /&gt;
          &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
          &amp;quot;image_url&amp;quot;: { &amp;quot;url&amp;quot;: (&amp;quot;data:image/png;base64,&amp;quot; + ($img | rtrimstr(&amp;quot;\n&amp;quot;))) }&lt;br /&gt;
        }&lt;br /&gt;
      ]&lt;br /&gt;
    }&lt;br /&gt;
  ]&lt;br /&gt;
}' &amp;gt; payload.json&lt;br /&gt;
&lt;br /&gt;
curl https://api-nothink.ailab.infocepo.com/v1/chat/completions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  --data-binary @payload.json&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import base64&lt;br /&gt;
import json&lt;br /&gt;
import requests&lt;br /&gt;
import os&lt;br /&gt;
&lt;br /&gt;
API_KEY = os.getenv(&amp;quot;OPENAI_API_KEY&amp;quot;)&lt;br /&gt;
MODEL = &amp;quot;ai-vision&amp;quot;&lt;br /&gt;
IMG_PATH = &amp;quot;/path/to/image.png&amp;quot;&lt;br /&gt;
API_URL = &amp;quot;https://api-nothink.ailab.infocepo.com/v1/chat/completions&amp;quot;&lt;br /&gt;
&lt;br /&gt;
with open(IMG_PATH, &amp;quot;rb&amp;quot;) as f:&lt;br /&gt;
    img_b64 = base64.b64encode(f.read()).decode(&amp;quot;utf-8&amp;quot;)&lt;br /&gt;
&lt;br /&gt;
payload = {&lt;br /&gt;
    &amp;quot;model&amp;quot;: MODEL,&lt;br /&gt;
    &amp;quot;messages&amp;quot;: [&lt;br /&gt;
        {&lt;br /&gt;
            &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;,&lt;br /&gt;
            &amp;quot;content&amp;quot;: [&lt;br /&gt;
                {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;Décris cette image.&amp;quot;},&lt;br /&gt;
                {&lt;br /&gt;
                    &amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;,&lt;br /&gt;
                    &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: f&amp;quot;data:image/png;base64,{img_b64}&amp;quot;}&lt;br /&gt;
                }&lt;br /&gt;
            ]&lt;br /&gt;
        }&lt;br /&gt;
    ]&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
headers = {&lt;br /&gt;
    &amp;quot;Authorization&amp;quot;: f&amp;quot;Bearer {API_KEY}&amp;quot;,&lt;br /&gt;
    &amp;quot;Content-Type&amp;quot;: &amp;quot;application/json&amp;quot;&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))&lt;br /&gt;
&lt;br /&gt;
if response.ok:&lt;br /&gt;
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))&lt;br /&gt;
else:&lt;br /&gt;
    print(f&amp;quot;Erreur {response.status_code}: {response.text}&amp;quot;)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API STT =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-audio2txt.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Modèle : &amp;lt;code&amp;gt;whisper-1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import requests&lt;br /&gt;
&lt;br /&gt;
OPENAI_API_KEY = 'sk-XXXXX'&lt;br /&gt;
&lt;br /&gt;
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'&lt;br /&gt;
headers = {&lt;br /&gt;
    'Authorization': f'Bearer {OPENAI_API_KEY}',&lt;br /&gt;
}&lt;br /&gt;
files = {&lt;br /&gt;
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),&lt;br /&gt;
    'model': (None, 'whisper-1')&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
response = requests.post(url, headers=headers, files=files)&lt;br /&gt;
print(response.json())&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
[ ! -f /tmp/test.ogg ] &amp;amp;&amp;amp; wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg&amp;quot; -O /tmp/test.ogg&lt;br /&gt;
&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -F model=&amp;quot;whisper-1&amp;quot; \&lt;br /&gt;
  -F file=&amp;quot;@/tmp/test.ogg&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
* Plusieurs formats audio sont acceptés.&lt;br /&gt;
* Le flux final est normalisé en '''16 kHz mono'''.&lt;br /&gt;
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.&lt;br /&gt;
&lt;br /&gt;
== UI ==&lt;br /&gt;
* [https://translate-rt.ailab.infocepo.com translate-rt]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API TTS =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-tts-omnivoice.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=sk-XXXXX&lt;br /&gt;
&lt;br /&gt;
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;model&amp;quot;: &amp;quot;gpt-4o-mini-tts&amp;quot;,&lt;br /&gt;
    &amp;quot;input&amp;quot;: &amp;quot;Bonjour, ceci est un test de synthèse vocale.&amp;quot;,&lt;br /&gt;
    &amp;quot;voice&amp;quot;: &amp;quot;coral&amp;quot;,&lt;br /&gt;
    &amp;quot;instructions&amp;quot;: &amp;quot;Speak in a cheerful and positive tone.&amp;quot;,&lt;br /&gt;
    &amp;quot;response_format&amp;quot;: &amp;quot;opus&amp;quot;&lt;br /&gt;
  }' | ffplay -i -&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text to Image =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-txt2image.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Clé API : &amp;lt;code&amp;gt;OPENAI_API_KEY=sk-...&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export OPENAI_API_KEY=EMPTY&lt;br /&gt;
&lt;br /&gt;
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer $OPENAI_API_KEY&amp;quot; \&lt;br /&gt;
  -d '{&lt;br /&gt;
    &amp;quot;prompt&amp;quot;: &amp;quot;a photo of a happy corgi puppy sitting and facing forward, studio light, longshot&amp;quot;,&lt;br /&gt;
    &amp;quot;n&amp;quot;: 1,&lt;br /&gt;
    &amp;quot;size&amp;quot;: &amp;quot;1024x1024&amp;quot;&lt;br /&gt;
  }'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Diarization =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
wget &amp;quot;https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3&amp;quot; -O /tmp/test.mp3&lt;br /&gt;
&lt;br /&gt;
curl -X POST &amp;quot;https://api-diarization.ailab.infocepo.com/upload-audio/&amp;quot; \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer token1&amp;quot; \&lt;br /&gt;
  -F &amp;quot;file=@/tmp/test.mp3&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Summary =&lt;br /&gt;
&lt;br /&gt;
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
text=&amp;quot;The tower is 324 metres tall and is one of the most recognizable monuments in the world.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
json_payload=$(jq -nc --arg text &amp;quot;$text&amp;quot; '{&amp;quot;text&amp;quot;: $text}')&lt;br /&gt;
&lt;br /&gt;
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \&lt;br /&gt;
  -H &amp;quot;Content-Type: application/json&amp;quot; \&lt;br /&gt;
  -d &amp;quot;$json_payload&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API Text Embeddings =&lt;br /&gt;
&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://api-embedding.ailab.infocepo.com/v1&amp;lt;/code&amp;gt;&lt;br /&gt;
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -k https://api-embedding.ailab.infocepo.com/v1/embeddings \&lt;br /&gt;
  -X POST \&lt;br /&gt;
  -d '{&amp;quot;model&amp;quot;:&amp;quot;bge-m3&amp;quot;,&amp;quot;input&amp;quot;:&amp;quot;What is Deep Learning?&amp;quot;}' \&lt;br /&gt;
  -H 'Content-Type: application/json'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= API DB Vectors (ChromaDB) =&lt;br /&gt;
&lt;br /&gt;
== Production ==&lt;br /&gt;
* URL : &amp;lt;code&amp;gt;https://chromadb.ailab.infocepo.com:wait-2026-12&amp;lt;/code&amp;gt;&lt;br /&gt;
* Token : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Lab ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12&lt;br /&gt;
export CHROMA_PORT=443&lt;br /&gt;
export CHROMA_TOKEN=XXXX&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple curl ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -v &amp;quot;${CHROMA_HOST}&amp;quot;/api/v1/collections \&lt;br /&gt;
  -H &amp;quot;Authorization: Bearer ${CHROMA_TOKEN}&amp;quot;&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple Python ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
import chromadb&lt;br /&gt;
from chromadb.config import Settings&lt;br /&gt;
&lt;br /&gt;
def chroma_http(host, port=80, token=None):&lt;br /&gt;
    return chromadb.HttpClient(&lt;br /&gt;
        host=host,&lt;br /&gt;
        port=port,&lt;br /&gt;
        ssl=host.startswith('https') or port == 443,&lt;br /&gt;
        settings=(&lt;br /&gt;
            Settings(&lt;br /&gt;
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',&lt;br /&gt;
                chroma_client_auth_credentials=token,&lt;br /&gt;
            ) if token else Settings()&lt;br /&gt;
        )&lt;br /&gt;
    )&lt;br /&gt;
&lt;br /&gt;
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)&lt;br /&gt;
collections = client.list_collections()&lt;br /&gt;
print(collections)&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Déployer sa propre instance ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
export nameSpace=your_namespace&lt;br /&gt;
domainRoot=ailab.infocepo.com&lt;br /&gt;
&lt;br /&gt;
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/&lt;br /&gt;
helm repo update&lt;br /&gt;
&lt;br /&gt;
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \&lt;br /&gt;
  --set chromadb.apiVersion=&amp;quot;0.4.24&amp;quot; \&lt;br /&gt;
  --set ingress.enabled=true \&lt;br /&gt;
  --set ingress.hosts[0].host=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot; \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].path=/ \&lt;br /&gt;
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \&lt;br /&gt;
  --set ingress.annotations.&amp;quot;cert-manager\.io/cluster-issuer&amp;quot;=letsencrypt-prod \&lt;br /&gt;
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \&lt;br /&gt;
  --set ingress.tls[0].hosts[0]=&amp;quot;${nameSpace}-chromadb.${domainRoot}&amp;quot;&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \&lt;br /&gt;
  -p '[{&amp;quot;op&amp;quot;:&amp;quot;add&amp;quot;,&amp;quot;path&amp;quot;:&amp;quot;/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size&amp;quot;,&amp;quot;value&amp;quot;:&amp;quot;0&amp;quot;}]'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Récupérer le token ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
kubectl --namespace ${nameSpace} get secret chromadb-auth \&lt;br /&gt;
  -o jsonpath=&amp;quot;{.data.token}&amp;quot; | base64 --decode &amp;amp;&amp;amp; echo&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Registry =&lt;br /&gt;
&lt;br /&gt;
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]&lt;br /&gt;
* Login : &amp;lt;code&amp;gt;user&amp;lt;/code&amp;gt;&lt;br /&gt;
* Password : &amp;lt;code&amp;gt;XXXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
curl -u &amp;quot;user:XXXXX&amp;quot; https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Exemple K8S ==&lt;br /&gt;
&amp;lt;pre&amp;gt;&lt;br /&gt;
deploymentName=&lt;br /&gt;
nameSpace=&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} create secret docker-registry pull-secret \&lt;br /&gt;
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \&lt;br /&gt;
  --docker-username=user \&lt;br /&gt;
  --docker-password=XXXXX \&lt;br /&gt;
  --docker-email=contact@example.com&lt;br /&gt;
&lt;br /&gt;
kubectl -n ${nameSpace} patch deployment ${deploymentName} \&lt;br /&gt;
  -p '{&amp;quot;spec&amp;quot;:{&amp;quot;template&amp;quot;:{&amp;quot;spec&amp;quot;:{&amp;quot;imagePullSecrets&amp;quot;:[{&amp;quot;name&amp;quot;:&amp;quot;pull-secret&amp;quot;}]}}}}'&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Stockage objet externe (S3) =&lt;br /&gt;
&lt;br /&gt;
* Endpoint : &amp;lt;code&amp;gt;https://s3.ailab.infocepo.com:wait-2026-09&amp;lt;/code&amp;gt;&lt;br /&gt;
* Access key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
* Secret key : &amp;lt;code&amp;gt;XXXX&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Un bucket nommé &amp;lt;code&amp;gt;ORG&amp;lt;/code&amp;gt; a été créé pour stocker des documents de démonstration.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= RAG optimisation =&lt;br /&gt;
&lt;br /&gt;
* Embeddings : &amp;lt;code&amp;gt;BAAI/bge-m3&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_size=1200&amp;lt;/code&amp;gt;&lt;br /&gt;
* &amp;lt;code&amp;gt;chunk_overlap=100&amp;lt;/code&amp;gt;&lt;br /&gt;
* LLM : &amp;lt;code&amp;gt;qwen3.6&amp;lt;/code&amp;gt;&lt;br /&gt;
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Workflow =&lt;br /&gt;
&lt;br /&gt;
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Environnements =&lt;br /&gt;
&lt;br /&gt;
== Hors production ==&lt;br /&gt;
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]&lt;br /&gt;
* Support : canal Mattermost Offre IA&lt;br /&gt;
* Le pseudo utilisateur doit respecter la convention interne&lt;br /&gt;
* Demander si besoin un accès Linux + Kubernetes&lt;br /&gt;
&lt;br /&gt;
== Production (best-effort) ==&lt;br /&gt;
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git&lt;br /&gt;
* Demander un namespace&lt;br /&gt;
* Lire la documentation de surveillance associée&lt;br /&gt;
&lt;br /&gt;
== Limites de l’infrastructure ==&lt;br /&gt;
* Les charges GPU sont intentionnellement limitées en journée.&lt;br /&gt;
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).&lt;br /&gt;
	&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Cloud Lab &amp;amp; projets d’audit =&lt;br /&gt;
&lt;br /&gt;
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]&lt;br /&gt;
&lt;br /&gt;
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.&lt;br /&gt;
&lt;br /&gt;
== Projet d’audit ==&lt;br /&gt;
; '''[[ServerDiff.sh]]'''&lt;br /&gt;
Script Bash d’audit permettant de :&lt;br /&gt;
* détecter les dérives de configuration,&lt;br /&gt;
* comparer plusieurs environnements,&lt;br /&gt;
* préparer un plan de migration ou de remédiation.&lt;br /&gt;
&lt;br /&gt;
== Exemple de migration cloud ==&lt;br /&gt;
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Tâche !! Description !! Durée (jours)&lt;br /&gt;
|-&lt;br /&gt;
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme d’architecture || Conception visuelle et documentation || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Vérification de stabilité || Premiers tests fonctionnels || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Étude d’automatisation || Identification des tâches répétitives || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Diagramme de migration || Illustration du processus || 1.0&lt;br /&gt;
|-&lt;br /&gt;
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Stabilisation || Validation de la reproductibilité || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Benchmark cloud || Comparaison vs legacy || 1.5&lt;br /&gt;
|-&lt;br /&gt;
| Réglage des temps d’arrêt || Calcul du downtime || 0.5&lt;br /&gt;
|-&lt;br /&gt;
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1&lt;br /&gt;
|-&lt;br /&gt;
! colspan=2 align=&amp;quot;right&amp;quot;| '''Total''' !! 15 jours.homme&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Vérifications de stabilité (HA minimale) ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Action !! Résultat attendu&lt;br /&gt;
|-&lt;br /&gt;
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds&lt;br /&gt;
|-&lt;br /&gt;
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Architecture web &amp;amp; bonnes pratiques =&lt;br /&gt;
&lt;br /&gt;
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes de conception :&lt;br /&gt;
&lt;br /&gt;
* privilégier une infrastructure '''simple, modulaire et flexible''',&lt;br /&gt;
* rapprocher le contenu du client (GDNS ou équivalent),&lt;br /&gt;
* utiliser des load balancers réseau (LVS, IPVS),&lt;br /&gt;
* comparer les coûts et éviter le '''vendor lock-in''',&lt;br /&gt;
* pour TLS :&lt;br /&gt;
** '''HAProxy''' pour les frontends rapides,&lt;br /&gt;
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),&lt;br /&gt;
* pour le cache :&lt;br /&gt;
** '''Varnish''', '''Apache Traffic Server''',&lt;br /&gt;
* favoriser les stacks open-source,&lt;br /&gt;
* utiliser files, buffers, queues et quotas pour lisser les pics.&lt;br /&gt;
&lt;br /&gt;
== Références ==&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]&lt;br /&gt;
* [https://github.com/systemdesign42/system-design System Design GitHub]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Comparatif des grandes plateformes cloud =&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt&lt;br /&gt;
|-&lt;br /&gt;
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python&lt;br /&gt;
|-&lt;br /&gt;
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API&lt;br /&gt;
|-&lt;br /&gt;
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A&lt;br /&gt;
|-&lt;br /&gt;
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Cette table sert de point de départ pour choisir la bonne stack selon :&lt;br /&gt;
* le niveau de contrôle souhaité,&lt;br /&gt;
* le contexte (on-prem, cloud public, HPC…),&lt;br /&gt;
* les outils d’automatisation existants.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Haute disponibilité, HPC &amp;amp; DevSecOps =&lt;br /&gt;
&lt;br /&gt;
== Haute disponibilité avec Corosync &amp;amp; Pacemaker ==&lt;br /&gt;
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]&lt;br /&gt;
&lt;br /&gt;
Principes :&lt;br /&gt;
* clusters multi-nœuds ou multi-sites,&lt;br /&gt;
* fencing via IPMI,&lt;br /&gt;
* provisioning PXE / NTP / DNS / TFTP,&lt;br /&gt;
* pour 2 nœuds : attention au split-brain,&lt;br /&gt;
* 3 nœuds ou plus recommandés en production.&lt;br /&gt;
&lt;br /&gt;
=== Ressources fréquentes ===&lt;br /&gt;
* multipath, LUNs, LVM, NFS,&lt;br /&gt;
* processus applicatifs,&lt;br /&gt;
* IP virtuelles, DNS, listeners réseau.&lt;br /&gt;
&lt;br /&gt;
== HPC ==&lt;br /&gt;
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]&lt;br /&gt;
&lt;br /&gt;
* orchestration de jobs (SLURM ou équivalent),&lt;br /&gt;
* stockage partagé haute performance,&lt;br /&gt;
* intégration possible avec des workloads IA.&lt;br /&gt;
&lt;br /&gt;
== DevSecOps ==&lt;br /&gt;
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]&lt;br /&gt;
&lt;br /&gt;
* CI/CD avec contrôles de sécurité intégrés,&lt;br /&gt;
* observabilité dès la conception,&lt;br /&gt;
* scans de vulnérabilité,&lt;br /&gt;
* gestion des secrets,&lt;br /&gt;
* policy-as-code.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= News &amp;amp; trends =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/@lev-selector/videos Top AI News]&lt;br /&gt;
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]&lt;br /&gt;
* [https://github.com/openai-translator/openai-translator OpenAI Translator]&lt;br /&gt;
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Formation &amp;amp; apprentissage =&lt;br /&gt;
&lt;br /&gt;
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]&lt;br /&gt;
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Liens cloud &amp;amp; IT utiles =&lt;br /&gt;
&lt;br /&gt;
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]&lt;br /&gt;
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]&lt;br /&gt;
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]&lt;br /&gt;
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia Cloud Wiki]&lt;br /&gt;
* [https://openapm.io OpenAPM]&lt;br /&gt;
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]&lt;br /&gt;
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]&lt;br /&gt;
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= Outils collaboratifs =&lt;br /&gt;
&lt;br /&gt;
== Dépôts de code ==&lt;br /&gt;
* [https://github.com/ynotopec GitHub ynotopec]&lt;br /&gt;
&lt;br /&gt;
== Base de connaissance ==&lt;br /&gt;
* ce wiki&lt;br /&gt;
&lt;br /&gt;
== Messagerie ==&lt;br /&gt;
* contact interne / support selon les projets&lt;br /&gt;
&lt;br /&gt;
== SSO ==&lt;br /&gt;
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]&lt;br /&gt;
&lt;br /&gt;
== MLflow ==&lt;br /&gt;
* [[MLFlow|MLFlow]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
= À propos &amp;amp; contributions =&lt;br /&gt;
&lt;br /&gt;
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.&lt;br /&gt;
&lt;br /&gt;
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.&lt;/div&gt;</summary>
		<author><name>Tcepo</name></author>
	</entry>
</feed>