巴基斯坦油气 HSE 场景下主权气隙 AI 栈的选型与成本设计
Sizing a sovereign, air-gapped AI stack for oil and gas health, safety and environment (HSE) in Pakistan
针对巴基斯坦油气运营商健康、安全与环境(HSE)部门的气隙部署需求,一份参考架构给出模型、硬件与成本方案:推理与智能体选用 753B 参数的 GLM 5.3(FP8 开放权重),时序预警用 amazon/chronos-2,视觉检测用 Roboflow/rf-detr-large,检索用 BAAI/bge-m3,OCR 用 PaddleOCR-VL 1.6。
This is the engineering summary of a full reference architecture. The paper, its object model as JSON and the model register are open: https://muhammadumar89.github.io/codeninja-research/sovereign-hse-pakistan/.
An oil and gas operator in Pakistan wants one thing from AI in its health, safety and environment department: a warning before the next incident, not a report after it. The data to do that already exists, spread across SAP EHS, SCADA and fire-and-gas historians, camera feeds and scanned investigation files. The constraint is just as clear. None of it may leave the operator's own infrastructure, and no third-party AI API may sit in the serving path.
Here is how that constraint turns into hardware, models and money.
1. Pick models by licence first
On an air-gapped platform you cannot call a hosted model, so every model must be self-hosted, and the licence decides whether the operator owns what it runs. Every pick lets the operator hold, run and fine-tune the weights inside its own boundary:
| Role | Model | Licence |
|---|---|---|
| Reasoning, cited answers, agents | GLM 5.3 open weights, 753B mixture-of-experts at FP8 | bespoke; purely internal use is exempt from its managed-service review |
| Time-series anomaly and early warning | amazon/chronos-2 | Apache-2.0 |
| Vision detection baseline | Roboflow/rf-detr-large (Nano to Large only) | Apache-2.0 |
| Tracking across frames | Roboflow trackers | Apache-2.0 |
| Multilingual retrieval (English, Urdu, Roman Urdu) | BAAI/bge-m3 | MIT |
| OCR of scanned permits and reports | PaddlePaddle/PaddleOCR-VL-1.6 | Apache-2.0 |
RF-DETR's larger checkpoints ship under a different platform licence, so the design stops at Large.
2. Size the central tier from the weights, not the brochure
Take the largest filed parameter count, multiply by bytes per parameter at the serving precision, then add a planning factor for the KV cache and activations so long incident histories fit:
weights = parameters x bytes per parameter
need = weights x 1.2 planning factor
nodes = ceil(need / (cards per node x memory per card))
GLM 5.3 is filed at 753B parameters. At FP8 that is 753 GB of weights and 904 GB with headroom, so one node of eight 141 GB cards (1,128 GB) holds it, leaving 375 GB beside the weights for KV cache: long incident histories and concurrent users.
3. Keep fast things at the edge
Detection and forecasting must keep up with cameras and sensors even if the link to the central tier drops. Detection runs on edge nodes inside the plant network, reusing the operator's NPU compute where it exists. Forecasting, OCR and embeddings run on a site inference server beside the historian: Chronos-2, PaddleOCR-VL 1.6 and BGE-M3 together weigh under 4 GB. Edge and site compute are sized by stream and decode load, not by model count.
4. One clock, one backbone, read-only adapters
- Every source enters through an adapter that only reads, tags provenance, maps to the object model once, is replayable and degrades honestly.
- Apache Kafka on KRaft orders events per equipment key, so a developing event is read in the order it happened.
- Chrony with a GNSS grandmaster gives sensors, cameras and servers one clock. Correlating SCADA with camera detections is meaningless if the timestamps drift.
5. What it costs, against the cloud
Three years, public prices, electricity at Pakistan's B3 industrial tariff:
| Option | Three-year cost (USD) |
|---|---|
| Own the hardware, with support and power | about 670,000 |
| Rent the same GPUs, AWS UAE region, three-year plan | 1.85 million |
| Rent the same GPUs, AWS UAE region, on demand | 2.82 million |
| Buy a closed frontier model by the token, 50 users | 1.04 to 2.95 million |
No hyperscaler runs a region inside Pakistan, so every rented option also moves the data abroad. The full workings and sources are in Appendix A of the paper.
6. The one hard dependency
141 GB HBM-class accelerators need a US export licence for Pakistan (Country Group D:4). Approved channels have delivered thousands of GPUs to Pakistani operators, and the design's first phase confirms the installed inventory before anything is bought, with a fallback to a mid-size model on existing hardware.
Reuse it
The object model ships as JSON in the repository in a format meant for import into an ontology platform, with every object's anchor system, properties, status vocabulary and links. Take it, change it, cite it.
Umar Bilal, Cofounder of CodeNinja. CodeNinja is a Middle Eastern American artificial intelligence lab that puts autonomy in physical operations.
来源:Google AI:DEV 作者专属(RSS) · dev.to