跳到正文
原文
Google AI:DEV 作者专属(RSS)· Martin D·· 9 小时前AI 评分21

香港 Databricks FSI Community Day 2026:用 Genie Ontology 扩展跨境流动性 Agent 推理

Scaling Cross-Border Liquidity Agent Inference with Genie Ontology Across Asia - Hong Kong Databricks FSI Community Day 2026

AI 导读

香港 Databricks FSI Community Day 2026 上的一场技术分享提出跨境流动性 Agent 推理架构,结合有界专家 Agent、Genie Ontology、Unity Catalog 与受治理检索,由监督 Agent 汇总证据并检测矛盾,但无权自行授予工具或执行资产划转。

正文

The Hong Kong Databricks FSI Community Day 2026 stands out as a highly unique, independent gathering happening directly within the Hong Kong Island waters. Operating away from typical convention centers, this exclusive, invitation-only event takes place entirely aboard a private boat traveling along the local ferry route. The forum serves as a dedicated working exchange for professionals operating at the intersection of complex data streams, financial markets, risk modeling, and institutional oversight.

To maintain absolute psychological and operational safety for its attendees, the organizers have stripped away traditional corporate hierarchies and product pitches in favor of open, critical peer challenges. There are no speaker names, titles, or recording devices permitted on board, ensuring that all field briefings focus strictly on executable expertise rather than corporate branding. Over thirty distinct technical proposals detail real-world financial architectures, handling everything from cross-border liquidity management and real-time streaming calculation paths to data isolation between entities in Hong Kong and Singapore. This community-driven event remains entirely independent of Databricks corporation, functioning instead as a private, expert-led ecosystem for practitioners navigating the realities of fragmented regional market structures.

Event Page:
https://vertexmacro.com/events/databricks_community_day_2026/index.html

Group Page:
https://usergroups.databricks.com/hong-kong-databricks-fsi-group/

Topic:
Scaling Cross-Border Liquidity Agent Inference with Genie Ontology Across Asia

Focus:
Optimizing Large-Scale AI Inference Infrastructure

Speaker Background:
Institutional AI infrastructure architect specializing in high-concurrency inference, distributed retrieval, agentic workflows, and governed financial data. The speaker designs platforms for Asian banks and trading firms, combining GPU efficiency, semantic control, legal-entity isolation, low-latency analytics, resilience, evidence, and model-risk governance.

Description:
Cross-border liquidity agents require more than a powerful language model. They must combine market data, treasury balances, settlement status, local regulations, funding costs, counterparty limits, and human authority while serving many users during volatile periods. Across Asia, the same term can have different meanings by jurisdiction, entity, currency, cut-off time, or settlement rail. Large-scale inference fails when throughput improves but semantic and legal boundaries disappear.

This session presents an inference architecture that combines bounded specialist agents, Genie Ontology, Unity Catalog, governed retrieval, and human-controlled workflows. One agent monitors a currency corridor. Another evaluates settlement delay. Another checks legal-entity positions and limits. A supervisor agent assembles evidence and detects contradictions, but it cannot grant itself tools, retrieve inaccessible data, or execute material asset movement.

Genie Ontology provides a business-aware context layer using governed semantics, domains, metric views, authoritative-source rules, and inferred organizational context. It maps fragmented local concepts into certified terms such as executable liquidity, trapped cash, final settlement, restricted balance, FX conversion premium, and stress-adjusted capacity. Context remains subject to Unity Catalog permissions. An agent cannot use an ontology snippet derived from an asset it is not entitled to access.

The inference path is decomposed into admission control, identity resolution, semantic planning, retrieval, SQL or analytical execution, reranking, model inference, tool calls, output validation, evidence assembly, and human disposition. Every stage carries legal entity, jurisdiction, currency corridor, purpose, model, agent version, ontology version, and policy context. Interactive questions receive strict latency and tool budgets. Long-running research moves to asynchronous workers with durable checkpoints and progress reporting.

GPU memory management determines realizable capacity. The session examines model weights, KV cache, activations, allocator fragmentation, context growth, host-to-device transfer, and concurrent tool waits. Validated techniques include quantization, paged attention, prefix reuse, context compression, tensor parallelism, model routing, and cache eviction. Financial accuracy, citation quality, numerical behavior, and multilingual performance are measured before an optimization reaches production.

Continuous batching groups compatible requests while balancing throughput, time-to-first-token, inter-token latency, deadlines, and fairness. Authority context is part of compatibility. Private HK and SG retrieval context is never combined or reused merely because prompts are similar. Cache keys include entity, jurisdiction, user or service identity, ontology version, data version, prompt policy, and model version. Shared public context may be reused globally; positions, balances, clients, and settlement evidence remain scoped.

Distributed retrieval is partitioned by entity, geography, domain, and authority. Metadata filters and permissions apply before or alongside semantic search so unauthorized candidates never enter the prompt. Retrieval observability covers index freshness, embedding version, document version, deletion propagation, source authority, residency, citation coverage, and denied candidates. Structured financial calculations use certified SQL or functions rather than asking the model to invent arithmetic.

During an Asia liquidity shock, hundreds of users ask for corridor status and alternatives. Admission control reserves capacity for treasury, risk, and incident queries. Continuous batching improves utilization, while fairness prevents one market from exhausting shared accelerators. Agents retrieve current balances, settlement status, approved regulatory context, and stress measures. The evidence panel shows route options, assumptions, expected premium, slippage, stale sources, unresolved conflicts, and why each source was permitted.

The human-in-the-loop boundary is explicit. Agents can investigate, compare hypotheses, and recommend. They cannot directly settle material funds, bypass limits, or invoke a critical asset transfer without an approved workflow. Human supervisors review evidence, select the route, apply limits, and accept accountability. Dual approval, tool allowlists, transaction limits, emergency revocation, and immutable audit records protect the final action.

Deployment uses an automated readiness protocol. Type, schema, semantic, idempotency, and permission tests run first. A shadow branch receives permitted read-only traffic and compares answers, calculations, citations, latency, cost, and policy decisions against production. Stress testing adds high concurrency, stale sources, model timeout, retrieval degradation, schema drift, and simulated liquidity shocks. Canary release limits users, entities, corridors, tools, and notional exposure.

Observability connects user experience to infrastructure and governance. Metrics include time-to-first-token, total latency, queue age, tokens per second, batch occupancy, GPU memory, cache hit rate, retrieval recall, tool latency, source freshness, ontology conflicts, policy denials, human override, cost per investigation, and realized versus expected routing outcome. Traces preserve reproducibility without retaining prohibited prompt content.

The final design treats inference as a regulated distributed service. Scale is valuable only when every recommendation remains permission-aware, semantically consistent, technically observable, financially bounded, and reviewable by an authorized human before capital moves across an Asian border.

Audience Takeaways:
Participants receive an agentic inference architecture, Genie Ontology semantic model, entity-safe batching and caching strategy, distributed-retrieval blueprint, GPU optimization framework, evidence-panel design, and shadow-release protocol. They will learn how to scale Asian liquidity intelligence while preserving permissions, semantic consistency, human control, auditability, resilience, and safe cross-border operations.

来源:Google AI:DEV 作者专属(RSS) · dev.to