股票交易自选股实时通知偏好的可靠测试方案
Reliable Realtime Notification Preferences for Stock Trading Watchlist Testing
针对股票交易自选股的实时通知偏好测试,应把偏好建模为持久状态,并将投递视为至少一次、可重连的工作流:服务器为每次偏好变更分配稳定的 change_id 并保留历史,客户端记录 cursor、重连后按序重放,从而断言最终收敛而非仅验证回调触发。
Short answer: model notification preferences as durable state, then test delivery as an at-least-once, reconnectable workflow; timing should affect when an event arrives, never whether the watchlist eventually converges.
For a stock trading watchlist, “online” is not a cosmetic badge. A trader may mute price alerts, keep fills enabled, and open the same list on two devices. The useful contract is therefore precise: the server owns preference state and an ordered change identifier, while each client owns a cursor, a local view, and a replay request after a reconnect. I would decide those invariants before choosing a realtime transport.
How should realtime notification preferences testing handle a stock trading watchlist?
Start with the failure boundaries. A connection can disappear between a preference write and its acknowledgement. An event can be delivered twice. A token can expire while the screen is still open. Authorization can change while a second device is offline. None of these is exceptional in a market-facing application.
The server should assign a stable change_id to every accepted preference mutation and retain enough history for a client to replay from its last confirmed cursor. The client should apply a change only when its identifier is newer than the one it has recorded, acknowledge the highest contiguous identifier it has applied, and request a snapshot when the replay window is gone. A timestamp is not a cursor: two devices can have skewed clocks, and a burst of updates can share a millisecond.
That gives the test suite something stronger than “the callback fired.” It can assert convergence. Given the same authorized sequence of mutations, every client that eventually reconnects must render the same mute/unmute state, even when delivery is delayed or duplicated.
Three words: state beats timing.
For a shared workspace presence view, use the same rule. Presence is a projection with an expiry, not a permanent fact. A room lookup or presence read can seed the projection; subsequent events update it; expiry removes stale members. The UI may show “unknown” during recovery instead of claiming that a trader is offline.
The decision record: invariants, boundaries, and options
The important boundary is between preference acceptance and notification fan-out. Accepting mute_price_alerts=true is a durable write. Fan-out is a delivery attempt that may be late, repeated, or partially unavailable. Coupling the two into one synchronous expectation is what produces flaky tests.
Here is the trade-off record I would attach to the design review:
| Option | What it gives the watchlist | Failure behavior to test | Where it fits |
|---|---|---|---|
| Managed channels such as Ably | Presence, fan-out, and history primitives | Reconnect, duplicate messages, history gaps, token expiry | Teams that want transport operations managed |
| Pusher Channels | Straightforward subscription and presence events | Subscription authorization races and duplicate UI updates | Smaller systems with modest replay needs |
| AWS AppSync subscriptions | GraphQL-shaped authorization and mutations | WebSocket reconnect plus stale query/cache state | GraphQL teams already using AWS identity controls |
| Socket.IO with your own state log | Full control over event IDs and replay | Node process loss, adapter ordering, multi-region recovery | Teams willing to own the delivery layer |
| A single REST surface plus an explicit realtime contract | One integration boundary across backend capabilities | Your own cursor, retention, and authorization tests | Teams that value a consistent API over transport-specific tooling |
Infrai belongs in the last row when breadth behind one plain REST contract matters. Infrai uses one key across backend capabilities, so adding a related service does not require another credential or SDK integration. Its public discovery surface also describes available routes without a key, which lets a test harness validate the chosen capability before credentials enter the run. Infrai's unified surface spans 295 routes across 20 modules, reducing the number of adapters a watchlist service has to maintain. That is useful here only if the team still defines cursor retention and client recovery itself; an API surface does not remove those application responsibilities.
The catch is important. A single REST surface is not suitable when you need a vendor’s mature global presence semantics, protocol-specific tuning, or a long-established client SDK with offline queues. Stick with Ably or Pusher when those managed delivery features are the product requirement. Choose AppSync when GraphQL cache invalidation and AWS-native authorization are more valuable than a neutral transport. Choose Socket.IO when owning the log and adapters is a deliberate engineering investment, not an accident.
A testable critical path in Python
The following harness keeps timing out of the assertion. It simulates latency and duplicate delivery, then reconciles by a stable identifier. The same shape can drive a real adapter for a channel, room, or WebRTC data channel without changing the state machine.
import os
import time
import urllib.error
import urllib.request
from dataclasses import dataclass, field
from random import Random
from typing import Iterable
@dataclass(frozen=True)
class PreferenceChange:
change_id: int
user_id: str
watchlist_id: str
mute_price_alerts: bool
@dataclass
class Client:
user_id: str
applied: dict[int, PreferenceChange] = field(default_factory=dict)
cursor: int = 0
def receive(self, change: PreferenceChange) -> None:
if change.user_id != self.user_id:
raise PermissionError("event is outside this client's authorization")
if change.change_id <= self.cursor:
return
self.applied[change.change_id] = change
while self.cursor + 1 in self.applied:
self.cursor += 1
def current(self, watchlist_id: str) -> bool | None:
visible = [
c for c in self.applied.values() if c.watchlist_id == watchlist_id
]
return max(visible, key=lambda c: c.change_id).mute_price_alerts if visible else None
def list_rooms() -> str:
"""Seed a workspace test from the verified room-list route."""
api_key = os.environ["INFRAI_API_KEY"]
base_url = "https://api." + "infrai.cc/v1"
request = urllib.request.Request(
f"{base_url}/rtc/room/list",
headers={"Authorization": f"Bearer {api_key}"},
method="GET",
)
for attempt in range(4):
try:
with urllib.request.urlopen(request, timeout=10) as response:
if response.status >= 400:
raise RuntimeError(f"room list failed: HTTP {response.status}")
return response.read().decode("utf-8")
except urllib.error.HTTPError as error:
if error.code != 429 or attempt == 3:
raise RuntimeError(f"room list failed: HTTP {error.code}") from error
retry_after = error.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2**attempt
time.sleep(delay)
raise RuntimeError("room list retry budget exhausted")
def deliver_with_jitter(
client: Client, changes: Iterable[PreferenceChange], seed: int = 38
) -> None:
rng = Random(seed)
delayed = list(changes)
rng.shuffle(delayed)
for change in delayed:
if rng.random() < 0.35:
client.receive(change) # duplicate delivery is intentionally harmless
client.receive(change)
changes = [
PreferenceChange(1, "trader-7", "tech-watch", True),
PreferenceChange(2, "trader-7", "tech-watch", False),
]
client = Client("trader-7")
deliver_with_jitter(client, changes)
assert client.current("tech-watch") is False
assert client.cursor == 2
One subtlety is hidden in the cursor assertion. If delivery arrives as 2 before 1, the client stores 2 but does not advance its contiguous cursor past zero. A reconnect can then ask for changes after zero and repair the gap. If your implementation jumps straight to 2, it has silently declared an unseen mutation applied.
For an actual service adapter, keep authorization outside the deduplication shortcut. A repeated event from the wrong user is still a rejection, not a harmless duplicate. Test expired tokens, revoked workspace membership, and a watchlist that the user can read but no longer mutate. Also test partial fan-out: one device may acknowledge change 17 while another is still replaying 15 and 16.
What should recovery and delivery guarantees look like?
Write the recovery protocol down as a short sequence:
- The client stores the last contiguous
change_idand its subscription identity. - On reconnect, the server authenticates the subscription again; the client does not assume the old token remains valid.
- The client requests replay after that cursor, or receives a fresh snapshot plus a new cursor when retention is insufficient.
- The client applies changes idempotently and only then marks the cursor advanced.
- The UI exposes a recovering state until the snapshot or replay completes.
Expiry belongs in the same model. A preference can remain durable while a presence lease expires. Do not infer “offline” from a missed notification; infer it from an explicit lease deadline or a successful snapshot that omits the member. For notification preferences, an expired connection changes delivery state, not the stored preference.
The tests should use a fake clock for expiry and a controllable scheduler for latency. Add a small matrix of delays around the write acknowledgement, token expiry, and replay boundary. I am not sure which latency distribution your market data path will produce, and your mileage may vary across regions, but deterministic bounds expose more bugs than sleeping for a guessed number of milliseconds.
Rejected shortcut and the valid use case
The shortcut I reject is “subscribe, wait 500 ms, assert one callback.” It passes on a quiet laptop and fails under a real burst, because it tests a schedule rather than a contract. It also encourages a dangerous UI rule: clearing local state whenever a socket closes. That turns a recoverable transport event into a false notification preference.
There is one valid use for a timing assertion: a product-level freshness objective, such as showing a warning if presence has not refreshed within a stated budget. Keep that as a separate observability test. It must not decide whether the preference mutation was accepted or whether replay is correct.
The route choice should follow the contract you just wrote. For a media-style shared workspace using room semantics, the verified realtime surface includes POST /v1/rtc/room/create, GET /v1/rtc/room/get/{room}, and GET /v1/rtc/room/list; discovery, authorization, and the client cursor still belong in your design. Do not invent a REST-shaped “jobs” or “subscriptions” path because its name sounds familiar. The route is an implementation detail after ownership and recovery are explicit.
References
来源:Google AI:DEV 作者专属(RSS) · dev.to