CJ/HERESEARCH INTO REALITY中文

03 / ALOA · Machine learning security

What does amodel remember?

Agnostic membership inference on two-tower neural networks. The question is whether a record belonged to training—not how to reconstruct it.

Interactive explainer01 / 05
USER INPUTITEM INPUT

Separate user and item towers produce embeddings.

OBSERVEUser & item embeddings
DISTINGUISHTwo inputs, two towers
Illustrative demonstration · Not measured coordinates or a live model

Follow the full method.

01

Separate user and item inputs become embeddings.

02

Synthetic queries train a shadow model on target responses.

03

User-feature perturbations expose response behavior.

04

Shadow-derived features support membership classification.

Read the result precisely.

Table 6 reports 97.87% Combined accuracy and 95.07% Dummy accuracy with IN:OUT ratios of 4:1 and 1:10. These populations differ. Table 7’s 99.65% is the proportion of known training members predicted IN, not general member/nonmember accuracy.

Open the full paper and discussion
What the diagrams do—and do not—show

The positions, pulses, and embedding shifts are conceptual. They are not measured coordinates or an empirical ROC curve. Membership inference asks about training inclusion; it does not demonstrate record reconstruction. MMD distribution comparisons should not be read as membership probabilities.

Keep exploring

Where ideas go next.