Main Page: Difference between revisions

From Essential
Jump to navigation Jump to search
 
(673 intermediate revisions by the same user not shown)
Line 1: Line 1:
Welcome to my experimental WIKI.
[[File:Infocepo-picture.png|thumb|right|Discover cloud and AI on infocepo.com]]
<br>
==[https://openai.com/ AI tools]==
==CLOUD LAB==
I want to share my [[LAB project]].<br>
<br>
[[file:Infocepo.drawio.png]]
==INFRA audit==
I made [[ServerDiff.sh]] script to audit servers.
You can track configuration drift.
You can check if your environments are the same.


==CLOUD migration example==
= infocepo.com – Cloud, AI & Labs =
*1.5 days: infra audit (82 clustered services) ([https://infocepo.com/wiki/index.php/ServerDiff.sh audit own tool])


*1.5 days: physical and virtual target CLOUD architecture diagram
Bienvenue sur le portail '''infocepo.com'''.


*1.5 days: physical compliance of 2 CLOUD (6 hypervisors, 6TB memory)
Ce wiki documente l’écosystème '''Cloud, IA, automatisation et lab''' d’Infocepo.
Il s’adresse aux :


*1 days: installation of the 2 CLOUD
* administrateurs systèmes,
* ingénieurs cloud,
* développeurs,
* étudiants,
* curieux qui veulent apprendre en pratiquant.


*.5 day: stability check
L’objectif est simple : transformer la théorie en '''scripts réutilisables, schémas, architectures, APIs et laboratoires concrets'''.
{| style="border-spacing:0;width:18.12cm;"
 
|- style="background-color:#ffc000;border:0.05pt solid #000000;padding:0.049cm;"
__TOC__
| align=center style="color:#000000;" | '''ACTION'''
 
| align=center style="color:#000000;" | '''RESULT'''
----
| align=center style="color:#000000;" | '''OK/KO'''
 
= Accès rapide =
 
== Portail principal ==
* [https://infocepo.com infocepo.com]
 
== Assistant IA ==
* [https://chat.infocepo.com Chat assistant]
 
== Liste des pages du wiki ==
* [[Special:AllPages|Toutes les pages]]
 
== Vue d’ensemble ==
[[File:Ailab-architecture.png|thumb|'''Infra architecture overview''']]
 
= Démarrer rapidement =
 
== Parcours recommandés ==
 
; 1. Construire un assistant IA privé
* Déployer une stack type '''Hermes WebUI + Ollama + GPU'''
* Ajouter un modèle de chat et un modèle de résumé
* Brancher des données internes via '''RAG + embeddings'''
 
; 2. Lancer un lab cloud
* Créer un petit cluster Kubernetes, OpenStack ou bare-metal
* Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)
* Ajouter un service IA : transcription, résumé, chatbot, OCR…
 
; 3. Préparer un audit ou une migration
* Inventorier les serveurs avec '''ServerDiff.sh'''
* Concevoir l’architecture cible
* Automatiser la migration avec des scripts reproductibles
 
== Vue d’ensemble du contenu ==
* '''Guides IA & outils''' : assistants, modèles, évaluation, GPU, RAG
* '''Cloud & infrastructure''' : Kubernetes, OpenStack, HA, HPC, DevSecOps
* '''Labs & scripts''' : audit, migration, automatisation
* '''Comparatifs''' : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.
 
----
 
= Vision =
 
[[File:Automation-full-vs-humans.png|thumb|right|The world after automation]]
 
Le but à long terme est de construire un environnement où :
 
* les assistants IA privés accélèrent la production,
* les tâches répétitives sont automatisées,
* les déploiements sont industrialisés,
* l’infrastructure reste '''compréhensible, portable et réutilisable'''.
 
Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.
 
[[File:SUMMARY-DIAGRAM-7311e6b1-aede-4989-ade2-a42d1a6e0ff2.png|thumb|right|Main page summary]]
 
----
 
= Catalogue rapide des services =
 
{| class="wikitable"
|+ Services principaux
! Catégorie !! Service !! Rôle
|-
| API || [https://api-llm-custom.ailab.infocepo.com/docs LLM] || Modèles de chat, code, RAG, OCR
|-
| API || [https://api-audio2txt.ailab.infocepo.com/docs STT] || Transcription audio
|-
| API || [https://api-tts-omnivoice.ailab.infocepo.com/docs TTS] || Synthèse vocale
|-
| API || [https://github.com/ynotopec/api-realtime-ai realtime-ai] || Temps réel WebSocket / WebRTC
|-
| API || [https://api-llm-custom.ailab.infocepo.com/docs IMAGE2TXT] || OCR / VLM via endpoint dédié
|-
| API || [https://api-summary.ailab.infocepo.com:wait-2026-12/docs summary] || Résumé de textes longs
|-
| API || [https://api-embedding.ailab.infocepo.com/docs EMBEDDINGS] || Embeddings pour RAG
|-
| API || [https://chromadb.ailab.infocepo.com:wait-2026-12 ChromaDB] || Base vecteur
|-
| API || [https://api-txt2image.ailab.infocepo.com/docs TXT2IMAGE] || Génération d’images
|-
| API || [https://api-diarization.ailab.infocepo.com/docs diarization] || Segmentation locuteurs
|-
|-
| style="border:0.05pt solid #000000;padding:0.049cm;color:#000000;" | Activate maintenance for n/2-1 nodes or 1 node if 2 nodes.
| Observabilité || [https://grafana.ailab.infocepo.com:wait-2026-12 monitoring] || Dashboards techniques
| style="border:0.05pt solid #000000;padding:0.049cm;color:#000000;" | All resources are started.
| style="background-color:#d8e4bc;border:0.05pt solid #000000;padding:0.049cm;color:#000000;" |  
|-
|-
| style="border:0.05pt solid #000000;padding:0.049cm;color:#000000;" | Un-maintenance all nodes. Power off n/2-1 nodes or 1 node if 2 nodes, different from the previous test.
| Observabilité || [https://uptime-kuma.ailab.infocepo.com:wait-2026-12/status/ai status] || Disponibilité des services
| style="border:0.05pt solid #000000;padding:0.049cm;color:#000000;" | All resources are started.
| style="background-color:#d8e4bc;border:0.05pt solid #000000;padding:0.049cm;color:#000000;" |  
|-
|-
| style="border:0.05pt solid #000000;padding:0.049cm;color:#000000;" | Power off simultaneous all nodes. Power on simultaneous all nodes.
| Observabilité || [https://web-stat.c1.ailab.infocepo.com:wait-2026-12 web-stat] || Statistiques web
| style="border:0.05pt solid #000000;padding:0.049cm;color:#000000;" | All resources are started.
| style="background-color:#d8e4bc;border:0.05pt solid #000000;padding:0.049cm;color:#000000;" |  
|-
|-
| Observabilité || [https://api.ailab.infocepo.com:wait-2026-09/ui LLM-stat] || Vue API / usage
|-
| Outils || [https://datalab.ailab.infocepo.com:wait-2026-12 dataLab] || Environnement de travail hors-production
|-
| Outils || [https://translate-rt.ailab.infocepo.com realtime translation] || Traduction
|-
| Outils || [https://demos.ailab.infocepo.com Demos] || Démonstrateurs
|}
|}
*1.5 days: CLOUD automation study


*1.5 days: 6 templates (2 CLOUD, 2 OS, 8 environments, 2 versions)
----
 
= Sélection AI & architecture par couche =
 
Le tableau de bord [https://ai.arch.infocepo.com '''Sélection AI'''] présente l'architecture par couche des modèles IA choisies :
 
{| class="wikitable"
! Couche !! Modèle !! Alias(s) !! Taille
|-
| rowspan="7" style="text-align:center; vertical-align:middle; background:#f0f0f0;"| '''HYBRID-CLOUD'''
| '''Qwen3.6-35B''' || flash, default, thinking, vision || Option 1, 35B MoE / 3B actif
|-
| '''Qwen3.8-Flash-Next''' || flash, default, thinking, vision, max || Option 2, 125B MoE, 51B n-grames internes, 6B actif, 1M contexte
|-
| '''MiniMax-H3''' || video || 33B dense
|-
| '''BGE-M3''' || embedding || 568M
|-
| '''BGE-reranker-v2-m3''' || reranker || 568M
|-
| '''FLUX.2-klein-4B''' || image || 4B
|-
| '''GLiNER2.5-Multi-V1''' || privacy || 287M, mDeBERTa-v3-base — privacy filter: [https://api-llm-privacy-proxy-gliner2.infocepo.com api-llm-privacy-proxy-gliner2]
|-
| '''Public-Cloud''' || '''GLM-5.3-Flash''' || max || 320B MoE / 18B actif — 1M contexte, multimodal
|-
| rowspan="2" style="text-align:center; vertical-align:middle; background:#f0f0f0;"| '''LOCAL'''
| '''Qwen3.5-9B''' || flash, default, thinking, vision || Option 1, 9B dense
|-
| '''Qwen3.8-27B''' || flash, default, thinking, vision, max || Option 2, 27B dense
|}
 
== Frameworks ==
 
{| class="wikitable"
! Outil !! Usage
|-
| '''Hermes''' || Usage général, raisonnement
|-
| '''OpenCode''' || DevSecOps
|-
| '''Kilo Code''' || Développement pur
|}
 
== Mémoire ==
 
Pour la couche d'embedding vectoriel (RAG, mémoire sémantique) :
 
* '''BGE-M3''' — 568M paramètres, architecture XLM-RoBERTa. Multi-fonction (dense + sparse + multi-vector), 100+ langues, licence MIT. Idéal comme moteur de mémoire à échelle.
 
= Nouveautés =
 
== Nouveautés implémentés 05/10/2026 ==
* [https://github.com/ynotopec/infocepo-cleanup '''infocepo-cleanup'''] (2026-10-03) : Nettoyage reproductible du wiki Main_Page InfoCEPO (dedup par slug GitHub)
* [https://github.com/ynotopec/docker-build '''docker-build'''] (2026-10-01) : Docker-in-Docker build factory on Kubernetes — build, push, pull via K8s DinD pod
* [https://github.com/ynotopec/gouv-fr-skills '''gouv-fr-skills'''] (2026-10-01) : gouv-fr-skills
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2-5 '''api-llm-privacy-proxy-gliner2-5'''] (2026-09-29) : api-llm-privacy-proxy-gliner2-5
* [https://github.com/ynotopec/api-audio2txt '''api-audio2txt'''] (2026-09-25) : Fast audio-to-text transcription API with real-time Whisper model support
* [https://github.com/ynotopec/translate-rt '''translate-rt'''] (2026-09-24) : Real-time translation API — multilingual translation with low latency
* [https://github.com/ynotopec/api-txt2image '''api-txt2image'''] (2026-09-20) : Text-to-image generation API for AI-powered image creation
* [https://github.com/ynotopec/api-rag '''api-rag'''] (2026-09-19) : RAG API — semantic search over knowledge bases with retrieval-augmented generation
* [https://github.com/ynotopec/gouv-fr-code-package-mi '''gouv-fr-code-package-mi'''] (2026-09-19) : gouv-fr-code-package-mi
* [https://github.com/ynotopec/agent-saas '''agent-saas'''] (2026-09-17) : SaaS AI agent platform — deploy autonomous AI agents as a service
* [https://github.com/ynotopec/qwen36-sglang '''qwen36-sglang'''] (2026-09-15) : Qwen 3.6 deployment with SGLang for high-performance inference
* [https://github.com/ynotopec/omnivoice-tts '''omnivoice-tts'''] (2026-09-13) : omnivoice-tts
* [https://github.com/ynotopec/api-llm-privacy-proxy-gliner2 '''api-llm-privacy-proxy-gliner2'''] (2026-09-10) : GLiNER2-powered LLM privacy proxy for named entity detection and filtering
* [https://github.com/ynotopec/tricoteuses-k8s '''tricoteuses-k8s'''] (2026-09-08) : Kubernetes deployment for tricoteuses-juridique
* [https://github.com/ynotopec/dashboard-superset-mcp '''dashboard-superset-mcp'''] (2026-09-03) : Automate Apache Superset dashboard creation via MCP server 6.1.0+
* [https://github.com/ynotopec/superset-k8s '''superset-k8s'''] (2026-09-01) : superset-k8s
* [https://github.com/ynotopec/quality-gate '''quality-gate'''] (2026-08-28) : quality-gate
* [https://github.com/ynotopec/api-convert2md '''api-convert2md'''] (2026-08-26) : Document-to-Markdown conversion API — clean markdown from various formats
* [https://github.com/ynotopec/trafilatura-local '''trafilatura-local'''] (2026-08-23) : trafilatura-local
* [https://github.com/ynotopec/models-todo '''models-todo'''] (2026-08-19) : models-todo
* [https://github.com/ynotopec/minimax-music3 '''minimax-music3'''] (2026-08-18) : minimax-music3
* [https://github.com/ynotopec/hermes-img-gen-infocepo '''hermes-img-gen-infocepo'''] (2026-08-15) : hermes-img-gen-infocepo
* [https://github.com/ynotopec/api-llm-custom '''api-llm-custom'''] (2026-08-13) : Custom LLM proxy API — route requests to any language model backend
* [https://github.com/ynotopec/infocepo-infra-mcp '''infocepo-infra-mcp'''] (2026-08-12) : infocepo-infra-mcp
* [https://github.com/ynotopec/api-mcp-openai '''api-mcp-openai'''] (2026-08-11) : AI MCP OpenAI Integration
* [https://github.com/ynotopec/api-embedding '''api-embedding'''] (2026-08-07) : Text embedding API — convert text to vector embeddings for semantic search
* [https://github.com/ynotopec/api-reranker '''api-reranker'''] (2026-08-06) : Cross-encoder reranking API for improving search retrieval quality
 
=== Top tasks — infrastructure déployée ===
''Généré depuis l'infra réelle (kubectl + endpoints LLM) le 2026-10-03. Ne pas éditer à la main — regénérer après déploiement.''
 
==== Endpoints LLM (vérifiés) ====
{| class="wikitable"
! Endpoint !! Modèle réel !! URL
|-
| '''ai-thinking''' || qwen3.6 (thinking) || <code>https://api-llm-custom.ailab.infocepo.com</code>
|-
| '''ai-default''' || qwen3.6 (défaut) || <code>https://api-llm-custom.ailab.infocepo.com</code>
|-
| '''ai-flash''' || qwen3.6 (fast) || <code>https://api-llm-custom.ailab.infocepo.com</code>
|-
| '''ai-vision''' || qwen3.6 (vision/OCR) || <code>https://api-llm-custom.ailab.infocepo.com</code>
|-
| '''ai-embedding''' || bge-m3 || <code>https://api-embedding.ailab.infocepo.com</code>
|}
 
==== Services & infrastructure IA ====
* [https://agents-saas.ailab.infocepo.com '''Agents SaaS'''] : Déploiement multi-tenant d'agents Hermes
* [https://bge-m3-article.ailab.infocepo.com '''bge-m3 Article'''] : Recherche sémantique sur articles (bge-m3)
* [https://infra2.ailab.infocepo.com '''Infra Dashboard v2'''] : Tableau de bord infrastructure
* [https://keycloak.demo1.ailab.infocepo.com '''Keycloak'''] : SSO / authentification centralisée
* [https://mcp-proxy.ailab.infocepo.com '''MCP Proxy'''] : Passerelle MCP
* [https://onyxia.demo1.ailab.infocepo.com '''Onyxia'''] : Data science platform (catalogue de services)
* [https://rag.ailab.infocepo.com '''RAG OpenWebUI'''] : RAG intégré à Open WebUI
* [https://registry.demo1.ailab.infocepo.com '''Registry'''] : Registre d'images Docker local
* [https://tika.demo1.ailab.infocepo.com '''Tika'''] : Extraction de texte (Apache Tika)
 
==== Benchmarks & évaluation ====
* [https://benchmark.demo1.ailab.infocepo.com '''Benchmark Endpoints'''] : Comparatif de performance d'endpoints LLM
* [https://eval.ailab.infocepo.com '''Eval Dashboard'''] : Évaluation de modèles LLM
* [https://qpr.infocepo.com '''OpenRouter Q/P'''] : Classement qualité/prix des modèles OpenRouter
* [https://oss-efficiency.ailab.infocepo.com '''OSS Efficiency'''] : Efficacité des modèles open-source
* [https://privacy.ailab.infocepo.com '''Privacy Benchmark'''] : Benchmark de fuites de secrets (proxy LLM)
* [https://rag-bench.demo1.ailab.infocepo.com '''RAG Bench'''] : Benchmark de pipelines RAG
* [https://rag-eval.demo1.ailab.infocepo.com '''RAG Eval'''] : Évaluation qualité RAG
 
==== Veille automatisée ====
* [https://models.ailab.infocepo.com '''Model Tracker'''] : Suivi des modèles HuggingFace vs déploiements
* [https://news-aggregator.demo1.ailab.infocepo.com '''News Aggregator'''] : Agrégation de flux d'actualité IA
 
=== Veille GitHub (populaire) ===
* [https://github.com/huggingface/transformers '''transformers'''] (2026-10-03) : Framework pour modèles de ML - transformers, embeddings, NLP (166k⭐)
* [https://github.com/unslothai/unsloth '''unsloth'''] (2026-10-03) : Interface locale pour exécuter et entraîner LLMs et modèles diffusion - GGUF, MLX, Qwen3.8, DeepSeek (77k⭐)
* [https://github.com/BerriAI/litellm '''litellm'''] (2026-10-03) : Passerelle AI rapide - Gateway Rust + SDK Python. 100+ APIs LLM en une (60k⭐)
* [https://github.com/vllm-project/vllm '''vllm'''] (2026-09-29) : A high-throughput and memory-efficient inference and serving engine for LLMs (93k⭐)
* [https://github.com/OpenHands/OpenHands '''OpenHands'''] (2026-09-29) : AI-Driven Development: open-source AI agent platform (89k⭐)
* [https://github.com/LibreChat-AI/LibreChat '''LibreChat'''] (2026-09-29) : Enhanced ChatGPT Clone: Features Agents, MCP, Skills, DeepSeek, Anthropic, AWS, OpenAI, Responses API (45k⭐)
* [https://github.com/n8n-io/n8n '''n8n'''] (2026-09-27) : Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or use cloud. (206k⭐)
* [https://github.com/QwenLM/qwen-code '''qwen-code'''] (2026-09-27) : An open-source AI coding agent that lives in your terminal. (28k⭐)
* [https://github.com/vercel/ai '''ai'''] (2026-09-27) : The AI Toolkit for TypeScript. From the creators of Next.js, the AI SDK is a free open-source library for building AI-powered chatbots, assistants and more. (26k⭐)
* [https://github.com/thedaviddias/Front-End-Checklist '''Front-End-Checklist'''] (2026-10-05) : 🗂 The essential checklist for modern web development, for humans and AI agents (74k⭐)
* [https://github.com/openclaw/openclaw '''openclaw'''] (2026-10-05) : The AI that really does things. Any OS. Any Platform. The lobster way. 🦞 (391k⭐)
* [https://github.com/promptfoo/promptfoo '''promptfoo'''] (2026-10-05) : Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanni… (26k⭐)
=== Backlog / Veille Technologique ===
 
=== Agents IA & Orchestration ===
* [https://github.com/langchain-ai/langchain '''LangChain'''] : Framework pour applications basées sur les LLM. Le plus mature. (145k⭐).
* [https://github.com/langgenius/dify '''Dify'''] : Plateforme de développement d'applications IA (LLM Ops). Déployable en self-host. (154k⭐).
* [https://github.com/open-webui/open-webui '''open-webui'''] : Interface web IA (Ollama, OpenAI API, MCP). (151k⭐).
* [https://github.com/ollama/ollama '''ollama'''] : Get up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models. (180k⭐).
* [https://github.com/milvus-io/milvus '''milvus'''] : Base de données vectorielle haute performance. (46k⭐).
=== Audio & TTS ===
* [https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct '''Qwen3-Omni-30B-A3B-Instruct'''] — Modèle multimodal Qwen (audio + texte + image).
 
=== Génération & Édition d'Images ===
* [https://github.com/Comfy-Org/ComfyUI '''ComfyUI''' ] : GUI et backend diffusion model le plus modulaire. (131k⭐)
 
=== RAG & Traitement de Documents ===
 
* [https://github.com/paperless-ngx/paperless-ngx '''paperless-ngx'''] : Document management : scan, index, archive. ML, OCR. (45k⭐).
* [https://github.com/graphify-labs/graphify '''graphify'''] : Transforme codebase en graphe de connaissances interrogeable. (115k⭐).
* [https://github.com/HKUDS/LightRAG '''LightRAG'''] : RAG simple et rapide. (39k⭐).
* [https://github.com/microsoft/graphrag '''graphrag'''] : RAG modulaire basé sur les graphes. (35k⭐).
* [https://github.com/headroomlabs-ai/headroom '''headroom'''] : Compress tool outputs, logs, files et RAG chunks. 20-95% token savings. (69k⭐).
* [https://github.com/mem0ai/mem0 '''Mem0'''] : — Mémorie à long terme pour agents IA. (64k⭐).
 
=== APIs à Développer ===
* '''Classificateur IA''' — Classification de contenu.
* '''Résumé mutualisé''' — API de résumé de texte partagée.
* '''NER''' — Reconnaissance d'entités nommées.
* '''Compressor''' — Compression de contenu.
 
=== Infrastructure & Backend ===
* [https://github.com/vllm-project/semantic-router '''Semantic Router'''] — Routage sémantique de requêtes vLLM.
 
=== Outils Dev ===
* [https://github.com/Aider-AI/aider '''Aider'''] — Assistant de codage IA en ligne de commande.
* [https://github.com/continuedev/continue '''Continue'''] — Extension IDE IA (VS Code, JetBrains). (35k⭐)
 
* [https://github.com/NVIDIA/TensorRT-LLM '''TensorRT-LLM'''] (2026-09-30) : TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-... (14k stars)
* [https://github.com/lutzroeder/netron '''netron'''] (2026-09-30) : Visualizer for neural network, deep learning and machine learning models (33k stars)
* [https://github.com/BasedHardware/omi '''omi'''] (2026-09-30) : AI that sees your screen, listens to your conversations and tells you what to do (13k stars)
----
= Assistants IA & outils cloud =
 
== Assistants IA ==
 
; '''Assistants IA auto-hébergés'''
* [https://github.com/nesquena/hermes-webui '''Hermes WebUI'''] + [https://ollama.com '''Ollama'''] + GPU 
: Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.
 
== Développement, modèles & veille ==
 
; '''Découverte de modèles'''
* [https://huggingface.co/models '''Models Trending''']
 
; '''Évaluation & benchmarks'''
* [https://arena.ai/leaderboard/code '''Agentic Evaluation''']
 
; '''Outils de développement & fine-tuning'''
* [https://github.com/trending?since=weekly '''Project Trending''']
* [https://grok.com '''News search''']
 
== Matériel IA & GPU ==
* NVIDIA GH200
* DGX Spark
* [https://www.mouser.fr/ProductDetail/BittWare/RS-GQ-GC1-0109?qs=ST9lo4GX8V2eGrFMeVQmFw%3D%3D '''GROQ LLM accelerator''']
 
----
 
= API Realtime AI (DEV) =
 
'''Statut :''' environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.
 
== Configuration ==
{| class="wikitable"
! Variable !! Valeur
|-
| OPENAI_API_BASE || <code>wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1</code>
|-
| OPENAI_API_KEY || <code>sk-XXXXX</code>
|}
 
== Dépôt GitHub ==
* [https://github.com/ynotopec/api-realtime-ai '''ynotopec/api-realtime-ai''']
 
== Page de test ==
* <code>external-test/half-duplex.html</code> — annulation d’écho + mode half-duplex.
 
== Compatibilité ==
Remplacer l’URL OpenAI par <code>$OPENAI_API_BASE</code> pour tester compatibilité et performances.
 
----
 
= API LLM (OpenAI compatible) =
 
* URL de base : <code>https://api-llm-custom.ailab.infocepo.com/v1</code>
* Création du token : [https://llm-token.ailab.infocepo.com:wait-2026-12 OPENAI_API_KEY]
* Documentation : [https://api-llm-custom.ailab.infocepo.com/docs Documentation API]
 
== Liste des modèles ==
<pre>
curl -X GET \
  'https://api-llm-custom.ailab.infocepo.com/v1/models' \
  -H 'Authorization: Bearer sk-XXXXX' \
  -H 'accept: application/json' \
  | jq | sed -rn 's#^.*id.*: "(.*)".*$#* \1#p' | sort -u
</pre>
 
== Modèles ouverts & endpoints internes ==
 
''Dernière mise à jour : 2026-10-03''
 
Les modèles ci-dessous correspondent à des '''endpoints logiques''' exposés derrière une passerelle.
 
{| class="wikitable"
! Endpoint !! Description / usage principal
|-
| '''ai-thinking''' || '''qwen3.6 fp8''' – thinking
|-
| '''ai-fast''' || '''qwen3.6 fp8''' en mode '''fast''' – vision/OCR/ai-default
|-
| '''ai-embedding''' || '''bge-m3''' – recherche sémantique
|-
| '''ai-stt-next''' || '''nemotron-3.5-asr-streaming-0.6b''' – streaming transcription vocale multilingual
|-
| '''ai-stt''' || '''whisper3-turbo''' – transcription vocale multilingual
|-
| '''ai-tts''' || '''OmniVoice''' – TTS multilingual
|-
| '''ai-image''' || '''FLUX.2-klein''' – generate and edit image
|}
 
== Exemple bash ==
<pre>
export OPENAI_API_MODEL="ai-default"
export OPENAI_API_BASE="https://api-llm-custom.ailab.infocepo.com/v1"
export OPENAI_API_KEY="sk-XXXXX"
 
promptValue="Quel est ton nom ?"
jsonValue='{
  "model": "'${OPENAI_API_MODEL}'",
  "messages": [{"role": "user", "content": "'${promptValue}'"}],
  "temperature": 0
}'
 
curl ${OPENAI_API_BASE}/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -d "${jsonValue}" 2>/dev/null | jq '.choices[0].message.content'
</pre>
 
== Vue infra LLM ==
[[File:Litellm-proxy-mermaid-diagram-2024-03-24-205202.png|thumb|right]]
 
'''DEV (au choix)'''
* '''A.''' <code>LiteLLM → vLLM/SgLang</code> : tests perf / compatibilité
* '''B.''' <code>LiteLLM → Ollama</code> : simple, rapide à itérer
* '''C.''' <code>Ollama</code> direct : POC ultra-léger
 
'''DEV – modèle FR / résumé'''
* <code>LiteLLM → Ollama /v1</code>
 
'''PROD'''
* '''Standard :''' <code>LiteLLM → vLLM/SgLang</code>
* '''Pont DEV→PROD :''' <code>LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang</code>
 
'''Notes :'''
* '''LiteLLM''' = passerelle unique (clés, quotas, logs)
* '''vLLM/SgLang''' = performance / stabilité en charge
* '''Ollama''' = simplicité de prototypage
 
----
 
= API Image to Text =
 
* Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.
* Modèle recommandé : <code>ai-vision</code>
 
== Exemple bash ==
<pre>
OPENAI_API_KEY=sk-XXXXX
 
base64 -w0 "/path/to/image.png" > img.b64
 
jq -n --rawfile img img.b64 \
'{
  model: "ai-vision",
  messages: [
    {
      role: "user",
      content: [
        { "type": "text", "text": "Décris cette image." },
        {
          "type": "image_url",
          "image_url": { "url": ("data:image/png;base64," + ($img | rtrimstr("\n"))) }
        }
      ]
    }
  ]
}' > payload.json
 
curl https://api-llm-custom.ailab.infocepo.com/v1/chat/completions \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  --data-binary @payload.json
</pre>
 
== Exemple Python ==
<pre>
import base64
import json
import requests
import os
 
API_KEY = os.getenv("OPENAI_API_KEY")
MODEL = "ai-vision"
IMG_PATH = "/path/to/image.png"
API_URL = "https://api-llm-custom.ailab.infocepo.com/v1/chat/completions"
 
with open(IMG_PATH, "rb") as f:
    img_b64 = base64.b64encode(f.read()).decode("utf-8")
 
payload = {
    "model": MODEL,
    "messages": [
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "Décris cette image."},
                {
                    "type": "image_url",
                    "image_url": {"url": f"data:image/png;base64,{img_b64}"}
                }
            ]
        }
    ]
}
 
headers = {
    "Authorization": f"Bearer {API_KEY}",
    "Content-Type": "application/json"
}
 
response = requests.post(API_URL, headers=headers, data=json.dumps(payload))
 
if response.ok:
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))
else:
    print(f"Erreur {response.status_code}: {response.text}")
</pre>
 
----
 
= API STT =
 
* URL : <code>https://api-audio2txt.ailab.infocepo.com/v1</code>
* Clé : <code>OPENAI_API_KEY=sk-XXXXX</code>
* Modèle : <code>whisper-1</code>
* Documentation : [https://api-audio2txt.ailab.infocepo.com/docs API STT docs]
 
== Exemple Python ==
<pre>
import requests
 
OPENAI_API_KEY = 'sk-XXXXX'
 
url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'
headers = {
    'Authorization': f'Bearer {OPENAI_API_KEY}',
}
files = {
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),
    'model': (None, 'whisper-1')
}
 
response = requests.post(url, headers=headers, files=files)
print(response.json())
</pre>
 
== Exemple curl ==
<pre>
[ ! -f /tmp/test.ogg ] && wget "https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg" -O /tmp/test.ogg
 
export OPENAI_API_KEY=sk-XXXXX
 
curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -F model="whisper-1" \
  -F file="@/tmp/test.ogg"
</pre>
 
== Notes ==
* Plusieurs formats audio sont acceptés.
* Le flux final est normalisé en '''16 kHz mono'''.
* Pour une qualité optimale : privilégier '''OPUS 16 kHz mono'''.
 
== UI ==
* [https://translate-rt.ailab.infocepo.com translate-rt]
 
----
 
= API TTS =
 
* URL : <code>https://api-tts-omnivoice.ailab.infocepo.com/v1</code>
* Clé : <code>OPENAI_API_KEY=sk-XXXXX</code>
* Documentation : [https://api-tts-omnivoice.ailab.infocepo.com/docs API TTS docs]
 
== Exemple ==
<pre>
export OPENAI_API_KEY=sk-XXXXX
 
curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o-mini-tts",
    "input": "Bonjour, ceci est un test de synthèse vocale.",
    "voice": "coral",
    "instructions": "Speak in a cheerful and positive tone.",
    "response_format": "opus"
  }' | ffplay -i -
</pre>
 
----
 
= API Text to Image =
 
* URL : <code>https://api-txt2image.ailab.infocepo.com/v1</code>
* Clé API : <code>OPENAI_API_KEY=sk-...</code>
* Documentation : [https://api-txt2image.ailab.infocepo.com/docs API TXT2IMAGE docs]
 
== Exemple ==
<pre>
export OPENAI_API_KEY=EMPTY
 
curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -d '{
    "prompt": "a photo of a happy corgi puppy sitting and facing forward, studio light, longshot",
    "n": 1,
    "size": "1024x1024"
  }'
</pre>
 
----


*1 day: migration diagram
= API Diarization =
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png]]


*1.5 days: 138 lines of industrialization code for migration ([https://infocepo.com/wiki/index.php/MigrationApp.sh migration own code])
* Documentation : [https://api-diarization.ailab.infocepo.com/docs API Diarization docs]


*1.5 days: process stabilization
== Exemple ==
<pre>
wget "https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3" -O /tmp/test.mp3


*1.5 days: CLOUD benchmark vs old INFRA
curl -X POST "https://api-diarization.ailab.infocepo.com/upload-audio/" \
  -H "Authorization: Bearer token1" \
  -F "file=@/tmp/test.mp3"
</pre>


*.5 days: calibration of unavailability time per unit migration
----


*5 minutes (effective load): 82 VM (env, os, application_code, 2 IP)
= API Summary =


Total = 15 man-days
* Documentation : [https://api-summary.ailab.infocepo.com:wait-2026-12/docs API Summary docs]


==CLOUD improvement==
== Exemple ==
[[File:WebModelDiagram.drawio.png]]
<pre>
*Formalize your infrastructure as much as possible for more flexibility, low complexity and less technology lock-in.
text="The tower is 324 metres tall and is one of the most recognizable monuments in the world."
*Use a name server able to handle the position of your customers like GDNS.
 
*Use a minimal instance and use a network load balancer like LVS. Monitor the global load of your instances and add/delete dynamically as needed.
json_payload=$(jq -nc --arg text "$text" '{"text": $text}')
*Or, many providers have dynamic computing services. Compare the prices. But take care about the technology lock-in.
 
*Use a very efficient TLS decoder like the HAPROXY decoder.
curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \
*Use very fast http cache like VARNISH.
  -H "Content-Type: application/json" \
*Use a big cache for big files like ATS.
  -d "$json_payload"
*...
</pre>
*Use serverless service for standard runtimes like Java, Python and PHP. But beware of certain incompatibilities and a lack of consistency over time.
 
*...
----
*Each time you need dynamic computing power think about load balancing or native service from the providers (caution about providers services!)
 
*...
= API Text Embeddings =
*Try to use open source STACKs as much as possible.
 
*...
* URL : <code>https://api-embedding.ailab.infocepo.com/v1</code>
*Use cache for your databases like MEMCACHED
* Documentation : [https://api-embedding.ailab.infocepo.com/docs Documentation]
 
== Exemple ==
<pre>
curl https://api-embedding.ailab.infocepo.com/v1/embeddings \
  -X POST \
  -d '{"model":"bge-m3","input":"What is Deep Learning?"}' \
  -H 'Content-Type: application/json'
</pre>
 
----
 
= API DB Vectors (ChromaDB) =
 
== Production ==
* URL : <code>https://chromadb.ailab.infocepo.com:wait-2026-12</code>
* Token : <code>XXXXX</code>
 
== Lab ==
<pre>
export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12
export CHROMA_PORT=443
export CHROMA_TOKEN=XXXX
</pre>
 
== Exemple curl ==
<pre>
curl -v "${CHROMA_HOST}"/api/v1/collections \
  -H "Authorization: Bearer ${CHROMA_TOKEN}"
</pre>
 
== Exemple Python ==
<pre>
import chromadb
from chromadb.config import Settings
 
def chroma_http(host, port=80, token=None):
    return chromadb.HttpClient(
        host=host,
        port=port,
        ssl=host.startswith('https') or port == 443,
        settings=(
            Settings(
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',
                chroma_client_auth_credentials=token,
            ) if token else Settings()
        )
    )
 
client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)
collections = client.list_collections()
print(collections)
</pre>
 
== Déployer sa propre instance ==
<pre>
export nameSpace=your_namespace
domainRoot=ailab.infocepo.com
 
helm repo add chroma https://amikos-tech.github.io/chromadb-chart/
helm repo update
 
helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \
  --set chromadb.apiVersion="0.4.24" \
  --set ingress.enabled=true \
  --set ingress.hosts[0].host="${nameSpace}-chromadb.${domainRoot}" \
  --set ingress.hosts[0].paths[0].path=/ \
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \
  --set ingress.annotations."cert-manager\.io/cluster-issuer"=letsencrypt-prod \
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \
  --set ingress.tls[0].hosts[0]="${nameSpace}-chromadb.${domainRoot}"
 
kubectl -n ${nameSpace} patch ingress/chromadb --type=json \
  -p '[{"op":"add","path":"/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size","value":"0"}]'
</pre>
 
== Récupérer le token ==
<pre>
kubectl --namespace ${nameSpace} get secret chromadb-auth \
  -o jsonpath="{.data.token}" | base64 --decode && echo
</pre>
 
----
 
= Registry =
 
* URL : [https://registry.ailab.infocepo.com:wait-2026-09 registry.ailab.infocepo.com:wait-2026-09]
* Login : <code>user</code>
* Password : <code>XXXXX</code>
 
== Exemple ==
<pre>
curl -u "user:XXXXX" https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog
</pre>
 
== Exemple K8S ==
<pre>
deploymentName=
nameSpace=
 
kubectl -n ${nameSpace} create secret docker-registry pull-secret \
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \
  --docker-username=user \
  --docker-password=XXXXX \
  --docker-email=contact@example.com
 
kubectl -n ${nameSpace} patch deployment ${deploymentName} \
  -p '{"spec":{"template":{"spec":{"imagePullSecrets":[{"name":"pull-secret"}]}}}}'
</pre>
 
----
 
= Stockage objet externe (S3) =
 
* Endpoint : <code>https://s3.ailab.infocepo.com:wait-2026-09</code>
* Access key : <code>XXXX</code>
* Secret key : <code>XXXX</code>
 
Un bucket nommé <code>ORG</code> a été créé pour stocker des documents de démonstration.
 
----
 
= RAG optimisation =
 
* Embeddings : <code>BAAI/bge-m3</code>
* <code>chunk_size=1000</code>
* <code>chunk_overlap=100</code>
* LLM : <code>qwen3.6</code>
* Pour les PDF mixtes : '''PDF → image → OCR / VLM''' peut améliorer les résultats.
 
----
 
= Workflow =
 
Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.
 
----
 
= Environnements =
 
== Hors production ==
* Utiliser [https://datalab.ailab.infocepo.com:wait-2026-12 datalab]
* Support : canal Mattermost Offre IA
* Le pseudo utilisateur doit respecter la convention interne
* Demander si besoin un accès Linux + Kubernetes
 
== Production (best-effort) ==
* Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git
* Demander un namespace
* Lire la documentation de surveillance associée
 
== Limites de l’infrastructure ==
* Les charges GPU sont intentionnellement limitées en journée.
* Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).
----
 
= Cloud Lab & projets d’audit =
 
[[File:Infocepo.drawio.png|400px|Cloud Lab reference diagram]]
 
Le '''Cloud Lab''' fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.
 
== Projet d’audit ==
; '''[[ServerDiff.sh]]'''
Script Bash d’audit permettant de :
* détecter les dérives de configuration,
* comparer plusieurs environnements,
* préparer un plan de migration ou de remédiation.
 
== Exemple de migration cloud ==
[[File:Diagram-migration-ORACLE-KVM-v2.drawio.png|400px|Cloud migration diagram]]


==CLOUD vs HW==
{| class="wikitable"
{| class="wikitable"
|'''Function'''
! Tâche !! Description !! Durée (jours)
|'''KUBERNETES'''
|-
|'''OPENSTACK'''
| Audit infrastructure || 82 services, audit automatisé via '''ServerDiff.sh''' || 1.5
|'''AWS'''
|-
|'''Bare-metal'''
| Diagramme d’architecture || Conception visuelle et documentation || 1.5
|'''HPC'''
|'''CRM'''
|'''OVIRT'''
|-
|-
|DEPLOY
| Contrôles de conformité || 2 clouds, 6 hyperviseurs, 6 To RAM || 1.5
|HELM/ANSIBLE/SH
|TERRAFORM/ANSIBLE/SH/JUJU
|TERRAFORM/CLOUDFOUNDATION/ANSIBLE/JUJU
|ANSIBLE/SH
|XCAT/CLUSH
|ANSIBLE/SH
|ANSIBLE/PYTHON/SH
|-
|-
|BOOTSTRAP
| Installation plateforme cloud || Déploiement des environnements cibles || 1.0
|API/CLI
|PXE/API/CLI
|API/CLI
|PXE/IPMI
|PXE/IPMI
|PXE/IPMI
|PXE/API
|-
|-
|
| Vérification de stabilité || Premiers tests fonctionnels || 0.5
|
|
|
|
|
|
|
|-
|-
|Router
| Étude d’automatisation || Identification des tâches répétitives || 1.5
|API/CLI (kube-router)
|API/CLI (router/subnet)
|API/CLI (Route table/subnet)
|LINUX/OVS/external
|XCAT/external
|LINUX/external
|API
|-
|-
|Firewall
| Développement des templates || 6 templates, 8 environnements, 2 clouds / OS || 1.5
|INGRESS/EGRESS/ISTIO
|API/CLI (Security groups)
|API/CLI (Security group)
|LINUX (NFT)
|LINUX (NFT)
|LINUX (NFT)
|API
|-
|-
|Vlan
| Diagramme de migration || Illustration du processus || 1.0
|DANM
|API/CLI (VPC)
|API/CLI (VPC)
|OVS/LINUX/external
|XCAT/external
|LINUX/external
|API
|-
|-
|
| Écriture du code de migration || 138 lignes (voir '''MigrationApp.sh''') || 1.5
|
|
|
|
|
|
|
|-
|-
|Name server
| Stabilisation || Validation de la reproductibilité || 1.5
|coredns
|dns-nameserver
|Amazon Route 53
|GDNS
|XCAT
|LINUX/external
|API/external
|-
|-
|Load balancer
| Benchmark cloud || Comparaison vs legacy || 1.5
|kube-proxy/LVS(IPVS)
|LVS
|Network Load Balancer
|LVS
|SLURM
|Ldirectord
|
|-
|-
|Storage
| Réglage des temps d’arrêt || Calcul du downtime || 0.5
|many
|-
|SWIFT/CINDER/NOVA
| Chargement VM || 82 VMs : OS, code, 2 IP par VM || 0.1
|S3/EFS/FSX/EBS
|-
|OPENSTACK SWIFT/XFS/EXT4/RAID10
! colspan=2 align="right"| '''Total''' !! 15 jours.homme
|GPFS
|}
|SAN
 
|NFS/SAN
=== Vérifications de stabilité (HA minimale) ===
{| class="wikitable"
! Action !! Résultat attendu
|-
| Extinction d’un nœud || Tous les services redémarrent automatiquement sur les autres nœuds
|-
| Extinction / redémarrage simultané de tous les nœuds || Les services repartent correctement après reboot
|}
|}


==[https://landscape.cncf.io/?fullscreen=yes CLOUD REF]==
==[https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison CLOUD providers]==
==[https://global-internet-map-2021.telegeography.com/ CLOUD map]==
==[https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Infrastructure example]==
==IT salaries==
*[http://jobsearchtech.about.com/od/educationfortechcareers/tp/HighestCerts.htm Best IT certifications]
*[https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital/ FREELANCE]
*[http://www.journaldunet.com/solutions/emploi-rh/salaire-dans-l-informatique-hays/ IT]


==[https://access.redhat.com/downloads/content/package-browser REDHAT package browser]==
'''Autonomie testée''' : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.
==HA COROSYNC+PACEMAKER==
 
===Typical architecture===
----
 
= Architecture web & bonnes pratiques =
 
[[File:WebModelDiagram.drawio.png|400px|Reference web architecture]]
 
Principes de conception :
 
* privilégier une infrastructure '''simple, modulaire et flexible''',
* rapprocher le contenu du client (GDNS ou équivalent),
* utiliser des load balancers réseau (LVS, IPVS),
* comparer les coûts et éviter le '''vendor lock-in''',
* pour TLS :
** '''HAProxy''' pour les frontends rapides,
** '''Envoy''' pour les cas avancés (mTLS, HTTP/2/3),
* pour le cache :
** '''Varnish''', '''Apache Traffic Server''',
* favoriser les stacks open-source,
* utiliser files, buffers, queues et quotas pour lisser les pics.
 
== Références ==
* [https://wikitech.wikimedia.org/wiki/Wikimedia_infrastructure Wikimedia infrastructure]
* [https://github.com/systemdesign42/system-design System Design GitHub]
 
----
 
= Comparatif des grandes plateformes cloud =
 
{| class="wikitable"
! Fonctionnalité !! Kubernetes !! OpenStack !! AWS !! Bare-metal !! HPC !! CRM !! oVirt
|-
| '''Outils de déploiement''' || Helm, YAML, ArgoCD, Juju || Ansible, Terraform, Juju || CloudFormation, Terraform, Juju || Ansible, Shell || xCAT, Clush || Ansible, Shell || Ansible, Python
|-
| '''Méthode de bootstrap''' || API || API, PXE || API || PXE, IPMI || PXE, IPMI || PXE, IPMI || PXE, API
|-
| '''Contrôle routeur''' || Kube-router || Router/Subnet API || Route Table / Subnet API || Linux, OVS || xCAT || Linux || API
|-
| '''Contrôle firewall''' || Istio, NetworkPolicy || Security Groups API || Security Group API || Linux firewall || Linux firewall || Linux firewall || API
|-
| '''Virtualisation réseau''' || VLAN, VxLAN || VPC || VPC || OVS, Linux || xCAT || Linux || API
|-
| '''DNS''' || CoreDNS || DNS-Nameserver || Route 53 || GDNS || xCAT || Linux || API
|-
| '''Load balancer''' || Kube-proxy, LVS || LVS || Network Load Balancer || LVS || SLURM || Ldirectord || N/A
|-
| '''Stockage''' || Local, cloud, PVC || Swift, Cinder, Nova || S3, EFS, EBS, FSx || Swift, XFS, EXT4, RAID10 || GPFS || SAN || NFS, SAN
|}
 
Cette table sert de point de départ pour choisir la bonne stack selon :
* le niveau de contrôle souhaité,
* le contexte (on-prem, cloud public, HPC…),
* les outils d’automatisation existants.
 
----
 
= Haute disponibilité, HPC & DevSecOps =
 
== Haute disponibilité avec Corosync & Pacemaker ==
[[File:HA-REF.drawio.png|400px|HA cluster architecture]]
 
Principes :
* clusters multi-nœuds ou multi-sites,
* fencing via IPMI,
* provisioning PXE / NTP / DNS / TFTP,
* pour 2 nœuds : attention au split-brain,
* 3 nœuds ou plus recommandés en production.
 
=== Ressources fréquentes ===
* multipath, LUNs, LVM, NFS,
* processus applicatifs,
* IP virtuelles, DNS, listeners réseau.
 
== HPC ==
[[File:HPC.drawio.png|400px|Overview of an HPC cluster]]
 
* orchestration de jobs (SLURM ou équivalent),
* stockage partagé haute performance,
* intégration possible avec des workloads IA.
 
== DevSecOps ==
[[File:DSO-POC-V3.drawio.png|400px|DevSecOps reference design]]
 
* CI/CD avec contrôles de sécurité intégrés,
* observabilité dès la conception,
* scans de vulnérabilité,
* gestion des secrets,
* policy-as-code.
 
----
 
= News & trends =
 
* [https://www.youtube.com/@lev-selector/videos Top AI News]
* [https://betterprogramming.pub/color-your-captions-streamlining-live-transcriptions-with-diart-and-openais-whisper-6203350234ef Real-time transcription with Diart + Whisper]
* [https://github.com/openai-translator/openai-translator OpenAI Translator]
* [https://opensearch.org/docs/latest/search-plugins/conversational-search Opensearch with LLM]
 
----
 
= Formation & apprentissage =
 
* [https://www.youtube.com/watch?v=4Bdc55j80l8 Transformers Explained]
* Labs, scripts et retours d’expérience concrets dans le projet Cloud Lab
 
----
 
= Liens cloud & IT utiles =
 
* [https://cloud.google.com/free/docs/aws-azure-gcp-service-comparison Cloud Providers Compared]
* [https://global-internet-map-2021.telegeography.com/ Global Internet Topology Map]
* [https://landscape.cncf.io/?fullscreen=yes CNCF Official Landscape]
* [https://openapm.io OpenAPM]
* [https://access.redhat.com/downloads/content/package-browser Red Hat Package Browser]
* [https://www.silkhom.com/barometre-2021-des-tjm-dans-informatique-digital Baromètre TJM IT]
* [https://www.glassdoor.fr/salaire/Hays-Salaires-E10166.htm Indicateurs salariaux IT]
 
----
 
= Outils collaboratifs =
 
== Dépôts de code ==
* [https://github.com/ynotopec '''GitHub ynotopec''']
 
== Base de connaissance ==
* [https://infocepo.com '''ce wiki''']


*2 rooms
== Messagerie ==
*2 power supply
* contact interne / support selon les projets
*2FC / server (active/active) (SAN)
*2*10Gbit/s ethernet / server (active/passive, possible active/active if PXE on native VLAN 0)
*IPMI VLAN (for the fence)
*VLAN ADMIN which must be the native VLAN if BOOTSTRAP by PXE (admin, provisioning, heartbeat)
*USER VLAN (application services)
*NTP
*DNS+DHCP+PXE+TFTP+HTTP for auto-provisioning
*PROXY (for update or otherwise internal REPOSITORY)


*Choose between 2 or more node clusters.
== SSO ==
* [https://auth-lab.ailab.infocepo.com:wait-2026-12/auth Keycloak]


*For a 2-node architecture, you need a 2-node configuration on COROSYNC and make sure to configure a 10-second staggered closing for one of the nodes (otherwise, an unstable cluster results).
----


*Resources are stateless.
= À propos & contributions =


For DB resources it is necessary to provide 4GB per base in general.
Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.
For CPU resources, as a rule there are no big requirements. Tip, for time-critical compressions, use PZSTD.


===Typical service pattern===
Ce wiki a vocation à rester un '''laboratoire vivant''' pour l’IA, le cloud et l’automatisation.
*MULTIPATH
*LUN
*LVM (LVM resource)
*FS (FS resource)
*NFS (FS resource)
*USER
*IP (IP resource)
*DNS name
*PROCESS (PROCESS resource)
*LISTENER (LISTENER resource)

Latest revision as of 09:19, 7 October 2026

Discover cloud and AI on infocepo.com

infocepo.com – Cloud, AI & Labs

Bienvenue sur le portail infocepo.com.

Ce wiki documente l’écosystème Cloud, IA, automatisation et lab d’Infocepo. Il s’adresse aux :

  • administrateurs systèmes,
  • ingénieurs cloud,
  • développeurs,
  • étudiants,
  • curieux qui veulent apprendre en pratiquant.

L’objectif est simple : transformer la théorie en scripts réutilisables, schémas, architectures, APIs et laboratoires concrets.


Accès rapide

Portail principal

Assistant IA

Liste des pages du wiki

Vue d’ensemble

Infra architecture overview

Démarrer rapidement

Parcours recommandés

1. Construire un assistant IA privé
  • Déployer une stack type Hermes WebUI + Ollama + GPU
  • Ajouter un modèle de chat et un modèle de résumé
  • Brancher des données internes via RAG + embeddings
2. Lancer un lab cloud
  • Créer un petit cluster Kubernetes, OpenStack ou bare-metal
  • Mettre en place un pipeline de déploiement (Helm, Ansible, Terraform…)
  • Ajouter un service IA : transcription, résumé, chatbot, OCR…
3. Préparer un audit ou une migration
  • Inventorier les serveurs avec ServerDiff.sh
  • Concevoir l’architecture cible
  • Automatiser la migration avec des scripts reproductibles

Vue d’ensemble du contenu

  • Guides IA & outils : assistants, modèles, évaluation, GPU, RAG
  • Cloud & infrastructure : Kubernetes, OpenStack, HA, HPC, DevSecOps
  • Labs & scripts : audit, migration, automatisation
  • Comparatifs : Kubernetes vs OpenStack vs AWS vs bare-metal, etc.

Vision

The world after automation

Le but à long terme est de construire un environnement où :

  • les assistants IA privés accélèrent la production,
  • les tâches répétitives sont automatisées,
  • les déploiements sont industrialisés,
  • l’infrastructure reste compréhensible, portable et réutilisable.

Chaque sortie IA est mesurée en temps réel, les modèles faibles sont remplacés automatiquement, et les tâches sans valeur sont coupées. L'autonomie vient de mécanismes techniques qui fonctionnent quand personne ne regarde — pas de règles écrites dans un doc.

Main page summary

Catalogue rapide des services

Services principaux
Catégorie Service Rôle
API LLM Modèles de chat, code, RAG, OCR
API STT Transcription audio
API TTS Synthèse vocale
API realtime-ai Temps réel WebSocket / WebRTC
API IMAGE2TXT OCR / VLM via endpoint dédié
API summary Résumé de textes longs
API EMBEDDINGS Embeddings pour RAG
API ChromaDB Base vecteur
API TXT2IMAGE Génération d’images
API diarization Segmentation locuteurs
Observabilité monitoring Dashboards techniques
Observabilité status Disponibilité des services
Observabilité web-stat Statistiques web
Observabilité LLM-stat Vue API / usage
Outils dataLab Environnement de travail hors-production
Outils realtime translation Traduction
Outils Demos Démonstrateurs

Sélection AI & architecture par couche

Le tableau de bord Sélection AI présente l'architecture par couche des modèles IA choisies :

Couche Modèle Alias(s) Taille
HYBRID-CLOUD Qwen3.6-35B flash, default, thinking, vision Option 1, 35B MoE / 3B actif
Qwen3.8-Flash-Next flash, default, thinking, vision, max Option 2, 125B MoE, 51B n-grames internes, 6B actif, 1M contexte
MiniMax-H3 video 33B dense
BGE-M3 embedding 568M
BGE-reranker-v2-m3 reranker 568M
FLUX.2-klein-4B image 4B
GLiNER2.5-Multi-V1 privacy 287M, mDeBERTa-v3-base — privacy filter: api-llm-privacy-proxy-gliner2
Public-Cloud GLM-5.3-Flash max 320B MoE / 18B actif — 1M contexte, multimodal
LOCAL Qwen3.5-9B flash, default, thinking, vision Option 1, 9B dense
Qwen3.8-27B flash, default, thinking, vision, max Option 2, 27B dense

Frameworks

Outil Usage
Hermes Usage général, raisonnement
OpenCode DevSecOps
Kilo Code Développement pur

Mémoire

Pour la couche d'embedding vectoriel (RAG, mémoire sémantique) :

  • BGE-M3 — 568M paramètres, architecture XLM-RoBERTa. Multi-fonction (dense + sparse + multi-vector), 100+ langues, licence MIT. Idéal comme moteur de mémoire à échelle.

Nouveautés

Nouveautés implémentés 05/10/2026

  • infocepo-cleanup (2026-10-03) : Nettoyage reproductible du wiki Main_Page InfoCEPO (dedup par slug GitHub)
  • docker-build (2026-10-01) : Docker-in-Docker build factory on Kubernetes — build, push, pull via K8s DinD pod
  • gouv-fr-skills (2026-10-01) : gouv-fr-skills
  • api-llm-privacy-proxy-gliner2-5 (2026-09-29) : api-llm-privacy-proxy-gliner2-5
  • api-audio2txt (2026-09-25) : Fast audio-to-text transcription API with real-time Whisper model support
  • translate-rt (2026-09-24) : Real-time translation API — multilingual translation with low latency
  • api-txt2image (2026-09-20) : Text-to-image generation API for AI-powered image creation
  • api-rag (2026-09-19) : RAG API — semantic search over knowledge bases with retrieval-augmented generation
  • gouv-fr-code-package-mi (2026-09-19) : gouv-fr-code-package-mi
  • agent-saas (2026-09-17) : SaaS AI agent platform — deploy autonomous AI agents as a service
  • qwen36-sglang (2026-09-15) : Qwen 3.6 deployment with SGLang for high-performance inference
  • omnivoice-tts (2026-09-13) : omnivoice-tts
  • api-llm-privacy-proxy-gliner2 (2026-09-10) : GLiNER2-powered LLM privacy proxy for named entity detection and filtering
  • tricoteuses-k8s (2026-09-08) : Kubernetes deployment for tricoteuses-juridique
  • dashboard-superset-mcp (2026-09-03) : Automate Apache Superset dashboard creation via MCP server 6.1.0+
  • superset-k8s (2026-09-01) : superset-k8s
  • quality-gate (2026-08-28) : quality-gate
  • api-convert2md (2026-08-26) : Document-to-Markdown conversion API — clean markdown from various formats
  • trafilatura-local (2026-08-23) : trafilatura-local
  • models-todo (2026-08-19) : models-todo
  • minimax-music3 (2026-08-18) : minimax-music3
  • hermes-img-gen-infocepo (2026-08-15) : hermes-img-gen-infocepo
  • api-llm-custom (2026-08-13) : Custom LLM proxy API — route requests to any language model backend
  • infocepo-infra-mcp (2026-08-12) : infocepo-infra-mcp
  • api-mcp-openai (2026-08-11) : AI MCP OpenAI Integration
  • api-embedding (2026-08-07) : Text embedding API — convert text to vector embeddings for semantic search
  • api-reranker (2026-08-06) : Cross-encoder reranking API for improving search retrieval quality

Top tasks — infrastructure déployée

Généré depuis l'infra réelle (kubectl + endpoints LLM) le 2026-10-03. Ne pas éditer à la main — regénérer après déploiement.

Endpoints LLM (vérifiés)

Endpoint Modèle réel URL
ai-thinking qwen3.6 (thinking) https://api-llm-custom.ailab.infocepo.com
ai-default qwen3.6 (défaut) https://api-llm-custom.ailab.infocepo.com
ai-flash qwen3.6 (fast) https://api-llm-custom.ailab.infocepo.com
ai-vision qwen3.6 (vision/OCR) https://api-llm-custom.ailab.infocepo.com
ai-embedding bge-m3 https://api-embedding.ailab.infocepo.com

Services & infrastructure IA

Benchmarks & évaluation

Veille automatisée

Veille GitHub (populaire)

  • transformers (2026-10-03) : Framework pour modèles de ML - transformers, embeddings, NLP (166k⭐)
  • unsloth (2026-10-03) : Interface locale pour exécuter et entraîner LLMs et modèles diffusion - GGUF, MLX, Qwen3.8, DeepSeek (77k⭐)
  • litellm (2026-10-03) : Passerelle AI rapide - Gateway Rust + SDK Python. 100+ APIs LLM en une (60k⭐)
  • vllm (2026-09-29) : A high-throughput and memory-efficient inference and serving engine for LLMs (93k⭐)
  • OpenHands (2026-09-29) : AI-Driven Development: open-source AI agent platform (89k⭐)
  • LibreChat (2026-09-29) : Enhanced ChatGPT Clone: Features Agents, MCP, Skills, DeepSeek, Anthropic, AWS, OpenAI, Responses API (45k⭐)
  • n8n (2026-09-27) : Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or use cloud. (206k⭐)
  • qwen-code (2026-09-27) : An open-source AI coding agent that lives in your terminal. (28k⭐)
  • ai (2026-09-27) : The AI Toolkit for TypeScript. From the creators of Next.js, the AI SDK is a free open-source library for building AI-powered chatbots, assistants and more. (26k⭐)
  • Front-End-Checklist (2026-10-05) : 🗂 The essential checklist for modern web development, for humans and AI agents (74k⭐)
  • openclaw (2026-10-05) : The AI that really does things. Any OS. Any Platform. The lobster way. 🦞 (391k⭐)
  • promptfoo (2026-10-05) : Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanni… (26k⭐)

Backlog / Veille Technologique

Agents IA & Orchestration

  • LangChain : Framework pour applications basées sur les LLM. Le plus mature. (145k⭐).
  • Dify : Plateforme de développement d'applications IA (LLM Ops). Déployable en self-host. (154k⭐).
  • open-webui : Interface web IA (Ollama, OpenAI API, MCP). (151k⭐).
  • ollama : Get up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models. (180k⭐).
  • milvus : Base de données vectorielle haute performance. (46k⭐).

Audio & TTS

Génération & Édition d'Images

  • ComfyUI  : GUI et backend diffusion model le plus modulaire. (131k⭐)

RAG & Traitement de Documents

  • paperless-ngx : Document management : scan, index, archive. ML, OCR. (45k⭐).
  • graphify : Transforme codebase en graphe de connaissances interrogeable. (115k⭐).
  • LightRAG : RAG simple et rapide. (39k⭐).
  • graphrag : RAG modulaire basé sur les graphes. (35k⭐).
  • headroom : Compress tool outputs, logs, files et RAG chunks. 20-95% token savings. (69k⭐).
  • Mem0 : — Mémorie à long terme pour agents IA. (64k⭐).

APIs à Développer

  • Classificateur IA — Classification de contenu.
  • Résumé mutualisé — API de résumé de texte partagée.
  • NER — Reconnaissance d'entités nommées.
  • Compressor — Compression de contenu.

Infrastructure & Backend

Outils Dev

  • Aider — Assistant de codage IA en ligne de commande.
  • Continue — Extension IDE IA (VS Code, JetBrains). (35k⭐)
  • TensorRT-LLM (2026-09-30) : TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-... (14k stars)
  • netron (2026-09-30) : Visualizer for neural network, deep learning and machine learning models (33k stars)
  • omi (2026-09-30) : AI that sees your screen, listens to your conversations and tells you what to do (13k stars)

Assistants IA & outils cloud

Assistants IA

Assistants IA auto-hébergés
Stack typique pour assistant privé, API OpenAI-compatible et expérimentation locale.

Développement, modèles & veille

Découverte de modèles
Évaluation & benchmarks
Outils de développement & fine-tuning

Matériel IA & GPU


API Realtime AI (DEV)

Statut : environnement DEV, remplaçante prévue de l’API OpenAI pour les cas temps réel.

Configuration

Variable Valeur
OPENAI_API_BASE wss://api-realtime-ai.ailab.infocepo.com:wait-2026-12/v1
OPENAI_API_KEY sk-XXXXX

Dépôt GitHub

Page de test

  • external-test/half-duplex.html — annulation d’écho + mode half-duplex.

Compatibilité

Remplacer l’URL OpenAI par $OPENAI_API_BASE pour tester compatibilité et performances.


API LLM (OpenAI compatible)

Liste des modèles

curl -X GET \
  'https://api-llm-custom.ailab.infocepo.com/v1/models' \
  -H 'Authorization: Bearer sk-XXXXX' \
  -H 'accept: application/json' \
  | jq | sed -rn 's#^.*id.*: "(.*)".*$#* \1#p' | sort -u

Modèles ouverts & endpoints internes

Dernière mise à jour : 2026-10-03

Les modèles ci-dessous correspondent à des endpoints logiques exposés derrière une passerelle.

Endpoint Description / usage principal
ai-thinking qwen3.6 fp8 – thinking
ai-fast qwen3.6 fp8 en mode fast – vision/OCR/ai-default
ai-embedding bge-m3 – recherche sémantique
ai-stt-next nemotron-3.5-asr-streaming-0.6b – streaming transcription vocale multilingual
ai-stt whisper3-turbo – transcription vocale multilingual
ai-tts OmniVoice – TTS multilingual
ai-image FLUX.2-klein – generate and edit image

Exemple bash

export OPENAI_API_MODEL="ai-default"
export OPENAI_API_BASE="https://api-llm-custom.ailab.infocepo.com/v1"
export OPENAI_API_KEY="sk-XXXXX"

promptValue="Quel est ton nom ?"
jsonValue='{
  "model": "'${OPENAI_API_MODEL}'",
  "messages": [{"role": "user", "content": "'${promptValue}'"}],
  "temperature": 0
}'

curl ${OPENAI_API_BASE}/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -d "${jsonValue}" 2>/dev/null | jq '.choices[0].message.content'

Vue infra LLM

DEV (au choix)

  • A. LiteLLM → vLLM/SgLang : tests perf / compatibilité
  • B. LiteLLM → Ollama : simple, rapide à itérer
  • C. Ollama direct : POC ultra-léger

DEV – modèle FR / résumé

  • LiteLLM → Ollama /v1

PROD

  • Standard : LiteLLM → vLLM/SgLang
  • Pont DEV→PROD : LiteLLM (DEV) → LiteLLM (PROD) → vLLM/SgLang

Notes :

  • LiteLLM = passerelle unique (clés, quotas, logs)
  • vLLM/SgLang = performance / stabilité en charge
  • Ollama = simplicité de prototypage

API Image to Text

  • Utilise l’API LLM avec un endpoint adapté à l’OCR / VLM.
  • Modèle recommandé : ai-vision

Exemple bash

OPENAI_API_KEY=sk-XXXXX

base64 -w0 "/path/to/image.png" > img.b64

jq -n --rawfile img img.b64 \
'{
  model: "ai-vision",
  messages: [
    {
      role: "user",
      content: [
        { "type": "text", "text": "Décris cette image." },
        {
          "type": "image_url",
          "image_url": { "url": ("data:image/png;base64," + ($img | rtrimstr("\n"))) }
        }
      ]
    }
  ]
}' > payload.json

curl https://api-llm-custom.ailab.infocepo.com/v1/chat/completions \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  --data-binary @payload.json

Exemple Python

import base64
import json
import requests
import os

API_KEY = os.getenv("OPENAI_API_KEY")
MODEL = "ai-vision"
IMG_PATH = "/path/to/image.png"
API_URL = "https://api-llm-custom.ailab.infocepo.com/v1/chat/completions"

with open(IMG_PATH, "rb") as f:
    img_b64 = base64.b64encode(f.read()).decode("utf-8")

payload = {
    "model": MODEL,
    "messages": [
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "Décris cette image."},
                {
                    "type": "image_url",
                    "image_url": {"url": f"data:image/png;base64,{img_b64}"}
                }
            ]
        }
    ]
}

headers = {
    "Authorization": f"Bearer {API_KEY}",
    "Content-Type": "application/json"
}

response = requests.post(API_URL, headers=headers, data=json.dumps(payload))

if response.ok:
    print(json.dumps(response.json(), indent=2, ensure_ascii=False))
else:
    print(f"Erreur {response.status_code}: {response.text}")

API STT

Exemple Python

import requests

OPENAI_API_KEY = 'sk-XXXXX'

url = 'https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions'
headers = {
    'Authorization': f'Bearer {OPENAI_API_KEY}',
}
files = {
    'file': ('file.opus', open('/path/to/file.opus', 'rb')),
    'model': (None, 'whisper-1')
}

response = requests.post(url, headers=headers, files=files)
print(response.json())

Exemple curl

[ ! -f /tmp/test.ogg ] && wget "https://upload.wikimedia.org/wikipedia/commons/1/17/Fables_de_La_Fontaine_Livre_1_01.ogg" -O /tmp/test.ogg

export OPENAI_API_KEY=sk-XXXXX

curl https://api-audio2txt.ailab.infocepo.com/v1/audio/transcriptions \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -F model="whisper-1" \
  -F file="@/tmp/test.ogg"

Notes

  • Plusieurs formats audio sont acceptés.
  • Le flux final est normalisé en 16 kHz mono.
  • Pour une qualité optimale : privilégier OPUS 16 kHz mono.

UI


API TTS

Exemple

export OPENAI_API_KEY=sk-XXXXX

curl https://api-tts-omnivoice.ailab.infocepo.com/v1/audio/speech \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o-mini-tts",
    "input": "Bonjour, ceci est un test de synthèse vocale.",
    "voice": "coral",
    "instructions": "Speak in a cheerful and positive tone.",
    "response_format": "opus"
  }' | ffplay -i -

API Text to Image

Exemple

export OPENAI_API_KEY=EMPTY

curl https://api-txt2image.ailab.infocepo.com/v1/images/generations \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -d '{
    "prompt": "a photo of a happy corgi puppy sitting and facing forward, studio light, longshot",
    "n": 1,
    "size": "1024x1024"
  }'

API Diarization

Exemple

wget "https://upload.wikimedia.org/wikipedia/commons/6/60/Mike_Peters_on_Politics_and_Emotion_%28Interview_1984%29.mp3" -O /tmp/test.mp3

curl -X POST "https://api-diarization.ailab.infocepo.com/upload-audio/" \
  -H "Authorization: Bearer token1" \
  -F "file=@/tmp/test.mp3"

API Summary

Exemple

text="The tower is 324 metres tall and is one of the most recognizable monuments in the world."

json_payload=$(jq -nc --arg text "$text" '{"text": $text}')

curl -X POST https://api-summary.ailab.infocepo.com:wait-2026-12/summary/ \
  -H "Content-Type: application/json" \
  -d "$json_payload"

API Text Embeddings

Exemple

curl https://api-embedding.ailab.infocepo.com/v1/embeddings \
  -X POST \
  -d '{"model":"bge-m3","input":"What is Deep Learning?"}' \
  -H 'Content-Type: application/json'

API DB Vectors (ChromaDB)

Production

Lab

export CHROMA_HOST=https://chromadb.c1.ailab.infocepo.com:wait-2026-12
export CHROMA_PORT=443
export CHROMA_TOKEN=XXXX

Exemple curl

curl -v "${CHROMA_HOST}"/api/v1/collections \
  -H "Authorization: Bearer ${CHROMA_TOKEN}"

Exemple Python

import chromadb
from chromadb.config import Settings

def chroma_http(host, port=80, token=None):
    return chromadb.HttpClient(
        host=host,
        port=port,
        ssl=host.startswith('https') or port == 443,
        settings=(
            Settings(
                chroma_client_auth_provider='chromadb.auth.token.TokenAuthClientProvider',
                chroma_client_auth_credentials=token,
            ) if token else Settings()
        )
    )

client = chroma_http(CHROMA_HOST, CHROMA_PORT, CHROMA_TOKEN)
collections = client.list_collections()
print(collections)

Déployer sa propre instance

export nameSpace=your_namespace
domainRoot=ailab.infocepo.com

helm repo add chroma https://amikos-tech.github.io/chromadb-chart/
helm repo update

helm upgrade --install chromadb chroma/chromadb -n ${nameSpace} \
  --set chromadb.apiVersion="0.4.24" \
  --set ingress.enabled=true \
  --set ingress.hosts[0].host="${nameSpace}-chromadb.${domainRoot}" \
  --set ingress.hosts[0].paths[0].path=/ \
  --set ingress.hosts[0].paths[0].pathType=ImplementationSpecific \
  --set ingress.annotations."cert-manager\.io/cluster-issuer"=letsencrypt-prod \
  --set ingress.tls[0].secretName=${nameSpace}-chromadb.${domainRoot}-tls \
  --set ingress.tls[0].hosts[0]="${nameSpace}-chromadb.${domainRoot}"

kubectl -n ${nameSpace} patch ingress/chromadb --type=json \
  -p '[{"op":"add","path":"/metadata/annotations/nginx.ingress.kubernetes.io~1proxy-body-size","value":"0"}]'

Récupérer le token

kubectl --namespace ${nameSpace} get secret chromadb-auth \
  -o jsonpath="{.data.token}" | base64 --decode && echo

Registry

Exemple

curl -u "user:XXXXX" https://registry.ailab.infocepo.com:wait-2026-09/v2/_catalog

Exemple K8S

deploymentName=
nameSpace=

kubectl -n ${nameSpace} create secret docker-registry pull-secret \
  --docker-server=registry.ailab.infocepo.com:wait-2026-09 \
  --docker-username=user \
  --docker-password=XXXXX \
  --docker-email=contact@example.com

kubectl -n ${nameSpace} patch deployment ${deploymentName} \
  -p '{"spec":{"template":{"spec":{"imagePullSecrets":[{"name":"pull-secret"}]}}}}'

Stockage objet externe (S3)

Un bucket nommé ORG a été créé pour stocker des documents de démonstration.


RAG optimisation

  • Embeddings : BAAI/bge-m3
  • chunk_size=1000
  • chunk_overlap=100
  • LLM : qwen3.6
  • Pour les PDF mixtes : PDF → image → OCR / VLM peut améliorer les résultats.

Workflow

Autoroute directe : code → CI/CD → deploy → monitoring. Pas de comité, pas de responsable désigné — chaque étape se déclenche automatiquement si la précédente passe. Si ça rate, alerte directe sans passer par un humain sauf décision explicite de fallback.


Environnements

Hors production

  • Utiliser datalab
  • Support : canal Mattermost Offre IA
  • Le pseudo utilisateur doit respecter la convention interne
  • Demander si besoin un accès Linux + Kubernetes

Production (best-effort)

  • Publier le code applicatif, les secrets (format SOPS), le Dockerfile et le code infra (Helm ou manifests K8S) sur Git
  • Demander un namespace
  • Lire la documentation de surveillance associée

Limites de l’infrastructure

  • Les charges GPU sont intentionnellement limitées en journée.
  • Coût unitaire mesuré automatiquement (dashboards temps réel, pas Excel).

Cloud Lab & projets d’audit

Cloud Lab reference diagram

Le Cloud Lab fournit des scénarios reproductibles : audit d’infrastructure, migration cloud, automatisation, haute disponibilité.

Projet d’audit

ServerDiff.sh

Script Bash d’audit permettant de :

  • détecter les dérives de configuration,
  • comparer plusieurs environnements,
  • préparer un plan de migration ou de remédiation.

Exemple de migration cloud

Cloud migration diagram

Tâche Description Durée (jours)
Audit infrastructure 82 services, audit automatisé via ServerDiff.sh 1.5
Diagramme d’architecture Conception visuelle et documentation 1.5
Contrôles de conformité 2 clouds, 6 hyperviseurs, 6 To RAM 1.5
Installation plateforme cloud Déploiement des environnements cibles 1.0
Vérification de stabilité Premiers tests fonctionnels 0.5
Étude d’automatisation Identification des tâches répétitives 1.5
Développement des templates 6 templates, 8 environnements, 2 clouds / OS 1.5
Diagramme de migration Illustration du processus 1.0
Écriture du code de migration 138 lignes (voir MigrationApp.sh) 1.5
Stabilisation Validation de la reproductibilité 1.5
Benchmark cloud Comparaison vs legacy 1.5
Réglage des temps d’arrêt Calcul du downtime 0.5
Chargement VM 82 VMs : OS, code, 2 IP par VM 0.1
Total 15 jours.homme

Vérifications de stabilité (HA minimale)

Action Résultat attendu
Extinction d’un nœud Tous les services redémarrent automatiquement sur les autres nœuds
Extinction / redémarrage simultané de tous les nœuds Les services repartent correctement après reboot


Autonomie testée : vérifier que la migration fonctionne seule en tirant le repo et lançant le script, sans personne qui connaît le système de tête.


Architecture web & bonnes pratiques

Reference web architecture

Principes de conception :

  • privilégier une infrastructure simple, modulaire et flexible,
  • rapprocher le contenu du client (GDNS ou équivalent),
  • utiliser des load balancers réseau (LVS, IPVS),
  • comparer les coûts et éviter le vendor lock-in,
  • pour TLS :
    • HAProxy pour les frontends rapides,
    • Envoy pour les cas avancés (mTLS, HTTP/2/3),
  • pour le cache :
    • Varnish, Apache Traffic Server,
  • favoriser les stacks open-source,
  • utiliser files, buffers, queues et quotas pour lisser les pics.

Références


Comparatif des grandes plateformes cloud

Fonctionnalité Kubernetes OpenStack AWS Bare-metal HPC CRM oVirt
Outils de déploiement Helm, YAML, ArgoCD, Juju Ansible, Terraform, Juju CloudFormation, Terraform, Juju Ansible, Shell xCAT, Clush Ansible, Shell Ansible, Python
Méthode de bootstrap API API, PXE API PXE, IPMI PXE, IPMI PXE, IPMI PXE, API
Contrôle routeur Kube-router Router/Subnet API Route Table / Subnet API Linux, OVS xCAT Linux API
Contrôle firewall Istio, NetworkPolicy Security Groups API Security Group API Linux firewall Linux firewall Linux firewall API
Virtualisation réseau VLAN, VxLAN VPC VPC OVS, Linux xCAT Linux API
DNS CoreDNS DNS-Nameserver Route 53 GDNS xCAT Linux API
Load balancer Kube-proxy, LVS LVS Network Load Balancer LVS SLURM Ldirectord N/A
Stockage Local, cloud, PVC Swift, Cinder, Nova S3, EFS, EBS, FSx Swift, XFS, EXT4, RAID10 GPFS SAN NFS, SAN

Cette table sert de point de départ pour choisir la bonne stack selon :

  • le niveau de contrôle souhaité,
  • le contexte (on-prem, cloud public, HPC…),
  • les outils d’automatisation existants.

Haute disponibilité, HPC & DevSecOps

Haute disponibilité avec Corosync & Pacemaker

HA cluster architecture

Principes :

  • clusters multi-nœuds ou multi-sites,
  • fencing via IPMI,
  • provisioning PXE / NTP / DNS / TFTP,
  • pour 2 nœuds : attention au split-brain,
  • 3 nœuds ou plus recommandés en production.

Ressources fréquentes

  • multipath, LUNs, LVM, NFS,
  • processus applicatifs,
  • IP virtuelles, DNS, listeners réseau.

HPC

Overview of an HPC cluster

  • orchestration de jobs (SLURM ou équivalent),
  • stockage partagé haute performance,
  • intégration possible avec des workloads IA.

DevSecOps

DevSecOps reference design

  • CI/CD avec contrôles de sécurité intégrés,
  • observabilité dès la conception,
  • scans de vulnérabilité,
  • gestion des secrets,
  • policy-as-code.

News & trends


Formation & apprentissage


Liens cloud & IT utiles


Outils collaboratifs

Dépôts de code

Base de connaissance

Messagerie

  • contact interne / support selon les projets

SSO


À propos & contributions

Suggestions de corrections, améliorations de schémas, retours d’expérience ou nouveaux labs bienvenus.

Ce wiki a vocation à rester un laboratoire vivant pour l’IA, le cloud et l’automatisation.