Encrypted Logistic Regression Prediction

Overview

The secure-inference page evaluated a linear model on encrypted patient data. Here the prediction also needs the logistic sigmoid

\[\sigma(\eta) = \frac{1}{1 + e^{-\eta}}.\]

The sigmoid is not a polynomial, and BFV/BGV and CKKS support only addition and multiplication. We replace it with a Chebyshev polynomial approximation, so the whole prediction (linear predictor and sigmoid) is computed on encrypted values.

The setting

A hospital holds patient data. A researcher holds a trained logistic-regression model. The researcher wants to score patients without seeing their data; the hospital wants predictions without seeing the model coefficients. This is the two-party setting of secure inference with a nonlinear model.

Step 1: train in the clear

The training set of 500 patients ships with the package.

import numpy as np
from homomorphepy import load_json
from homomorphepy.examples.encrypted_regression import fit_logistic

fx = load_json("encrypted_regression")
age       = np.asarray(fx["age"], dtype=float)
biomarker = np.asarray(fx["biomarker"], dtype=float)
outcome   = np.asarray(fx["outcome"], dtype=float)

beta = fit_logistic(age, biomarker, outcome)
print("Coefficients (intercept, age, biomarker):", np.round(beta, 4))
Coefficients (intercept, age, biomarker): [-2.2671  0.0348  0.8991]

Step 2: encrypt patient data

The Chebyshev polynomial needs enough multiplicative depth. We use depth 8 and enable PKESchemeFeature.ADVANCEDSHE, which provides the polynomial-evaluation functions.

import openfhe
from homomorphepy import fhe_context, packed_codec, Ct

cc = fhe_context("CKKS",
                 multiplicative_depth=8,
                 scaling_mod_size=50,
                 batch_size=16,
                 features=[openfhe.PKESchemeFeature.ADVANCEDSHE])
keys = cc.KeyGen()
cc.EvalMultKeyGen(keys.secretKey)
codec = packed_codec(cc)

new_age = [45, 52, 60, 38, 70, 55, 48, 63,
           41, 57, 66, 44, 72, 50, 59, 35]
new_bm  = [-0.5,  0.3, 1.2, -1.0, 0.8, 0.1, -0.3, 1.5,
           -0.8,  0.6, 0.9, -0.4, 1.1, 0.0,  0.7, -1.2]

ct_age = Ct(cc.Encrypt(keys.publicKey, codec.encode(new_age)), cc.cc)
ct_bm  = Ct(cc.Encrypt(keys.publicKey, codec.encode(new_bm)),  cc.cc)

Step 3: evaluate the linear predictor (encrypted)

\(\eta = \beta_0 + \beta_1\, \text{age} + \beta_2\, \text{biomarker}\), all arithmetic on encrypted data:

ct_eta = ct_age * float(beta[1])
ct_eta = ct_eta + ct_bm * float(beta[2])
ct_eta = ct_eta + float(beta[0])

Step 4: apply the sigmoid (encrypted)

OpenFHE exposes Chebyshev approximations of common transcendental functions, including the logistic sigmoid, EvalLogistic(). The interval \([a, b]\) must cover the range of \(\eta\) for these patients. A degree-16 approximation fits within the depth set above; its error is measured in Step 5.

ct_prob = Ct(cc.cc.EvalLogistic(ct_eta.raw, -4.0, 4.0, 16), cc.cc)

Step 5: decrypt predictions

encrypted_probs = np.asarray(
    codec.decode(cc.Decrypt(ct_prob.raw, keys.secretKey), 16), dtype=float)

eta_clear       = beta[0] + beta[1] * np.asarray(new_age) + beta[2] * np.asarray(new_bm)
cleartext_probs = 1.0 / (1.0 + np.exp(-eta_clear))
max_err = float(np.max(np.abs(encrypted_probs - cleartext_probs)))
print(f"Max absolute error vs cleartext sigmoid: {max_err:.2e}")
Max absolute error vs cleartext sigmoid: 5.79e-06

The same run, in one call

from homomorphepy.examples import encrypted_regression

r = encrypted_regression.run()
r.max_error
5.79341590434268e-06

Summary

  1. Model coefficients are applied to encrypted data. The researcher never sees the patient data. The hospital decrypts the predictions but is not sent the coefficients, as in secure inference; see the threat-model caveat below.
  2. The sigmoid is evaluated on encrypted values. There is no intermediate decryption and no exchange between the parties during the computation.
  3. The CKKS error is small. The maximum absolute error against the cleartext sigmoid is 5.8e-06.

Limitations

  • Multiplicative depth: each multiplication uses one level of the depth set in the context. Deeper computations need larger parameters and more memory.
  • Approximation error: CKKS is approximate. For deployment, precision would need to be validated for the use case.
  • Performance: encrypted computation is orders of magnitude slower than cleartext.
  • Threat model: the model-extraction caveat of the secure-inference page applies here. The hospital holds the secret key and could recover the coefficients from predictions on chosen inputs. The sigmoid makes this harder than in the linear case but does not prevent it. A deployment would need further protection, such as output differential privacy, query limits, or threshold FHE.