SoxAIDocs
Security

Data Loss Prevention (DLP)

Scan every AI request for credentials, PII, and sensitive data before it reaches the upstream model.

Data Loss Prevention (DLP)

Overview

DLP runs in the relay hot path after channel selection and before the upstream request. It scans the JSON request body for sensitive content and applies a policy action — mask, block, or audit — before the request leaves SoxAI. Masking rewrites the body in place; the upstream receives a redacted version. Blocking rejects the request before it reaches any external service.

Relay Pipeline
  tokenAuth → quotaCheck → channelSelect
       ↓
  ✦ applyDLP          ← scans + rewrites/blocks here
       ↓
  applyRequestTransform
       ↓
  upstream (model provider)

Builtin Detectors

SoxAI ships 15 builtin detectors. All are enabled by default and can be suppressed per policy.

CodeDescriptionPost-validatorCategory
credit_cardVisa, Mastercard, Amex, DiscoverLuhn algorithmFinancial
cn_id_cardChinese national ID (18-digit)GB 11643 check digitPII
openai_api_keyOpenAI sk-... keys—Credentials
anthropic_api_keyAnthropic sk-ant-... keys—Credentials
aws_access_keyAWS AKIA... access key IDs—Credentials
aws_secret_key40-char AWS secret keys—Credentials
gcp_service_keyGCP service account JSON keys—Credentials
generic_jwteyJ... JWT tokens (any algorithm)—Credentials
pem_private_keyPEM-encoded RSA/EC/Ed25519 private keys—Credentials
us_ssnUS Social Security Numbers (NNN-NN-NNNN)—PII
emailEmail addresses—PII
cn_phoneChinese mobile numbers (11-digit 1xx)—PII
ipv4_internalPrivate IPv4 addressesRFC 1918 rangesPII
bitcoin_walletBitcoin mainnet addresses—Financial
mac_addressIEEE 802 MAC addresses—PII

Custom Detectors

Create tenant-private detectors in Console → Security → DLP → Detectors → New Detector.

Regex Detectors

Use RE2 syntax. Supply a confidence score (0–100) that flows into policy decisions. Optionally add a post-validator to reduce false positives.

Example — match an internal project code like PROJ-12345:

Pattern:   \bPROJ-\d{5}\b
Confidence: 90

Dictionary Detectors

Supply a newline-separated list of terms. The engine uses Aho-Corasick for O(n) multi-term matching. Useful for codenames, internal system identifiers, or customer lists.

Example terms:

project-phoenix
operation-nightfall
internal-codename-x

Policies

A policy binds one or more detectors to a scope and an action. Create policies in Console → Security → DLP → Policies → New Policy.

Actions

ActionWhat happens
maskMatched spans are replaced with [REDACTED_<detector_code>]. The rewritten body is forwarded to the upstream. The caller receives a normal response.
blockThe request is rejected before reaching the upstream. The caller receives a 403 error.
audit_onlyThe request passes unchanged. A finding is recorded for review.

Scope Bindings

A policy can target any combination of scopes. Multiple bindings use union semantics — if any scope matches, the policy applies.

ScopeDescription
TeamAll requests from members of a specific team
UserRequests from a specific user
TokenRequests authenticated with a specific API token

Minimum Confidence

Set a minimum confidence threshold (0–100) per policy. Findings below the threshold are ignored by that policy. This lets you enable a detector globally (e.g., email) but only act on high-confidence matches.

API Behavior

When a Request Is Blocked

HTTP status 403 Forbidden. Response body:

{
  "error": {
    "code": "dlp_blocked",
    "message": "Request blocked: credit card number detected in prompt.",
    "detector": "credit_card",
    "request_id": "req_01abc..."
  }
}

message describes which detector triggered. detector is the detector code (e.g. credit_card, openai_api_key). request_id correlates with gateway logs.

When a Request Is Masked

The request completes normally (200). The upstream receives the rewritten body. The caller is not notified that masking occurred — the response reflects what the model produced given the redacted input.

When Audit-Only

The request completes normally (200). No visible change to the caller or the upstream.

Findings & Audit

Every mask and block finding (and all audit_only findings) is encrypted and stored.

  • Encryption: AES-256-GCM. The matched text is never stored in plaintext.
  • Viewing findings: Console → Security → DLP → Findings. Findings show detector code, action taken, timestamp, and scope metadata. The matched text is stored encrypted and is not visible until decrypted.
  • Decrypting a finding: Only system_admin accounts can decrypt. Every decrypt triggers an immutable entry in the audit log. Navigate to a finding and click Decrypt — you will be prompted for step-up authentication.

See Also