Secure Model Inference on Encrypted Data

The use case

A diagnostic lab develops a proprietary scoring model for disease risk based on a panel of biomarkers. A hospital wants to score its patients but cannot send raw patient data to the lab in cleartext. The lab does not want to send its model coefficients to the hospital in cleartext either.

With FHE the hospital encrypts patient biomarkers, sends the encrypted values to the lab, the lab applies its weighted-sum scoring to them while they stay encrypted, and the hospital decrypts the scores. The biomarker values never appear in cleartext on the lab’s machine, and the lab’s coefficients are never sent to the hospital. These are the transport guarantees and they are what the pipeline below illustrates. They are not the whole story of deploying a model-as-a-service — see Threat model: model extraction from the hospital side below.

This is the two-party companion to the multi-site master/worker pattern (MLE, Cox). Unlike those, there is no aggregator and no consensus: one client (hospital) holds the secret key and queries one server (lab) which holds the model coefficients in cleartext. The homomorphic primitives all come from the openfhe-python binding.

Scenario: weighted biomarker score

The lab’s model computes a weighted score:

\[\text{score} = w_1 x_1 + w_2 x_2 + w_3 x_3 + w_4 x_4 + b\]

where \(x_i\) are biomarker values and \(w_i, b\) are proprietary coefficients.

Setup: hospital’s context and patient data

from homomorphepy import fhe_context, packed_codec, Ct

cc = fhe_context("CKKS", multiplicative_depth=2,
                 scaling_mod_size=50, batch_size=8)
keys = cc.KeyGen()
cc.EvalMultKeyGen(keys.secretKey)     # for encrypted * encrypted; unused below
codec = packed_codec(cc)

# 8 patients, each with 4 biomarker values. We pack each biomarker
# across patients: one encrypted value per biomarker, with the patient
# values in its slots, so one operation acts on every patient at once.
biomarker1 = [1.2, 0.8, 1.5, 0.3, 2.1, 0.9, 1.1, 1.8]
biomarker2 = [0.5, 1.1, 0.3, 0.8, 0.2, 1.4, 0.7, 0.6]
biomarker3 = [2.0, 1.5, 2.3, 1.0, 1.8, 2.1, 1.6, 2.5]
biomarker4 = [0.1, 0.4, 0.2, 0.6, 0.3, 0.1, 0.5, 0.2]

ct1, ct2, ct3, ct4 = (
    Ct(cc.Encrypt(keys.publicKey, codec.encode(b)), cc.cc)
    for b in (biomarker1, biomarker2, biomarker3, biomarker4)
)

Lab side: apply the model to encrypted data

The lab receives the encrypted biomarkers and applies its proprietary model — without ever seeing patient values.

# Lab's proprietary model weights (never shared with the hospital)
w = [0.35, -0.20, 0.50, 0.15]
b = 1.2

# Encrypted score = w1*x1 + w2*x2 + w3*x3 + w4*x4 + b
ct_score = ct1 * w[0] + ct2 * w[1] + ct3 * w[2] + ct4 * w[3] + b

Hospital side: decrypt the results

pt = cc.Decrypt(ct_score.raw, keys.secretKey)
scores = codec.decode(pt, 8)

def risk_band(s):
    return "HIGH" if s > 2.0 else "MODERATE" if s > 1.5 else "LOW"

import pandas as pd

pd.DataFrame({
    "patient": range(1, 9),
    "score":   [round(s, 3) for s in scores],
    "risk":    [risk_band(s) for s in scores],
})
patient score risk
0 1 2.535 HIGH
1 2 2.070 HIGH
2 3 2.845 HIGH
3 4 1.735 MODERATE
4 5 2.840 HIGH
5 6 2.300 HIGH
6 7 2.320 HIGH
7 8 2.990 HIGH

Verification

import numpy as np

cleartext = (np.array(w) @ np.array([biomarker1, biomarker2,
                                     biomarker3, biomarker4]) + b)
max_error = float(np.max(np.abs(np.array(scores) - cleartext)))
max_error
8.348877145181177e-14

CKKS gives essentially the same answer as cleartext, within floating-point precision.

assert max_error < 1e-6

The same run, in one call

from homomorphepy.examples import secure_inference

result = secure_inference.run()
result.max_error
1.2434497875801753e-13

Transport guarantees

The pipeline above delivers two concrete protections:

  1. Biomarker values never appear in cleartext outside the hospital. The lab’s view of the protocol consists of the encrypted values it received and the encrypted score it returned.
  2. The lab’s coefficients \(w\) and \(b\) are never sent to the hospital in cleartext. They are used only to construct the returned encrypted score inside the lab’s Python session.
Party Cleartext view
Hospital Patient biomarkers (local), decrypted scores
Lab Encrypted values only — no cleartext biomarker values, no cleartext scores

These are necessary conditions for any model-as-a-service deployment that does not trust the lab with cleartext patient data. They are not sufficient conditions, as the next section shows.

Threat model: model extraction from the hospital side

The hospital holds the secret key and decides what goes into the encrypted queries. Nothing in the FHE pipeline restricts the biomarker values the hospital encrypts. For a linear model with four biomarkers and a bias, the hospital can recover every coefficient with five queries by submitting the standard basis:

  • \(\mathbf{e}_0 = (0,0,0,0)\ \Rightarrow\ \text{score} = b\)
  • \(\mathbf{e}_1 = (1,0,0,0)\ \Rightarrow\ \text{score} = w_1 + b\)
  • \(\mathbf{e}_2 = (0,1,0,0)\ \Rightarrow\ \text{score} = w_2 + b\)
  • \(\mathbf{e}_3 = (0,0,1,0)\ \Rightarrow\ \text{score} = w_3 + b\)
  • \(\mathbf{e}_4 = (0,0,0,1)\ \Rightarrow\ \text{score} = w_4 + b\)

Subtracting the first score from each of the others recovers the four weights exactly. We can run this attack in the same Python session: wrap the lab’s scoring pipeline as a function that closes over \(w\) and \(b\) without revealing them, then pack the five probes across slots 1–5 of the four encrypted biomarker vectors.

# Lab pipeline wrapped as a function. Closes over `w` and `b`; the
# caller (hospital) never reads either.
def lab_score(ct_bio):
    return (ct_bio[0] * w[0] + ct_bio[1] * w[1]
            + ct_bio[2] * w[2] + ct_bio[3] * w[3] + b)

# Hospital crafts five probes packed across slots 1..5. Slot 1 is e_0
# (all zeros, probes b). Slot j+1 is e_j (a one in position j, probes
# w_j + b).
probe_bio = [
    [0, 1, 0, 0, 0, 0, 0, 0],
    [0, 0, 1, 0, 0, 0, 0, 0],
    [0, 0, 0, 1, 0, 0, 0, 0],
    [0, 0, 0, 0, 1, 0, 0, 0],
]
ct_probe = [Ct(cc.Encrypt(keys.publicKey, codec.encode(v)), cc.cc) for v in probe_bio]

probe_pt = cc.Decrypt(lab_score(ct_probe).raw, keys.secretKey)
probe_scores = np.asarray(codec.decode(probe_pt, 5))

b_hat = probe_scores[0]
w_hat = probe_scores[1:5] - b_hat

pd.DataFrame([[b] + w, [b_hat] + list(w_hat)],
             index=["true", "recovered"],
             columns=["b", "w1", "w2", "w3", "w4"]).round(6)
b w1 w2 w3 w4
true 1.2 0.35 -0.2 0.5 0.15
recovered 1.2 0.35 -0.2 0.5 0.15

The recovered coefficients match the lab’s true coefficients to CKKS precision. Five queries — one bias probe plus one per biomarker — are enough because the scoring function is linear in the biomarkers; a linear function of \(k\) inputs is fully specified by any \(k + 1\) affinely independent point evaluations. FHE does not impede this extraction: the hospital is the party that decrypts, the scoring function is the thing being released, and the standard basis is a legal query vector.

The same attack in one call:

attack = secure_inference.extract_model()
f"{attack['n_queries']} queries, max weight error {attack['max_weight_error']:.1e}"
'5 queries, max weight error 2.0e-13'

Layered defenses

A deployment that needs the scoring function to stay proprietary has to layer defenses on top of the FHE transport:

  • Output perturbation. Adding calibrated noise to the decrypted score — a differential-privacy mechanism — turns each query into a noisy observation rather than an exact readout. Extraction accuracy degrades with the noise variance and a query budget can be accounted for explicitly.
  • Query-budget accounting. Capping the number of queries a client can issue in a session bounds how much of the model can leak before the budget is exhausted. This composes with output perturbation rather than replacing it.
  • Threshold FHE. Splitting the secret key across several parties (see the Cox threshold page for the master/worker case) changes who can decrypt and under what audit trail. The hospital alone can no longer mount the adaptive-query attack; decryption requires cooperation from the other key-holders.
  • Non-linear scoring structure. A linear model is the trivial extraction case. Non-linear scorers (the logistic model of encrypted regression, polynomial compositions) resist closed-form basis-vector attacks, though approximation attacks from the model-extraction literature still apply.

The page above shows the transport mechanics end-to-end. It is a toy illustration in that the linear scorer plus hospital-held secret key is exactly the configuration where the five-query attack works; a production model-as-a-service deployment is the transport layer plus at least one of the defenses above.

Network protocol: serialization

In practice the hospital and lab are on different machines. All objects serialize for network transport:

import tempfile
from pathlib import Path

import openfhe

tdir = Path(tempfile.mkdtemp())
openfhe.SerializeToFile(str(tdir / "context.bin"), cc.cc, openfhe.BINARY)
openfhe.SerializeToFile(str(tdir / "pubkey.bin"), keys.publicKey, openfhe.BINARY)
openfhe.SerializeToFile(str(tdir / "patient_bm1.bin"), ct1.raw, openfhe.BINARY)

# Lab receives the serialized files
cc_lab, _ = openfhe.DeserializeCryptoContext(str(tdir / "context.bin"), openfhe.BINARY)
ct_lab, _ = openfhe.DeserializeCiphertext(str(tdir / "patient_bm1.bin"), openfhe.BINARY)

# Lab applies its weights to the deserialized encrypted value
ct_weighted = Ct(ct_lab, cc_lab) * 0.35

# Lab returns the result
openfhe.SerializeToFile(str(tdir / "weighted.bin"), ct_weighted.raw, openfhe.BINARY)

# Hospital receives, deserializes, decrypts
ct_recv, _ = openfhe.DeserializeCiphertext(str(tdir / "weighted.bin"), openfhe.BINARY)
codec.decode(cc.Decrypt(ct_recv, keys.secretKey), 8)
[0.4200000000001037,
 0.2799999999998495,
 0.5249999999999823,
 0.10500000000002324,
 0.7349999999999528,
 0.315000000000102,
 0.384999999999901,
 0.6299999999999977]