Hashed ElGamal under CCA1 and Hybrid Encryption

1. Overview

This lecture has two main parts. The first part finishes the discussion of public-key encryption by proving that Hashed ElGamal is secure against non-adaptive chosen-ciphertext attacks in the random-oracle model, under the particular Strong Diffie–Hellman assumption used in the course. The second part introduces hybrid encryption and key encapsulation mechanisms (KEMs), explains their efficiency and security definitions, gives Diffie–Hellman and Hashed ElGamal KEMs as examples, and shows how a CCA-secure KEM can be combined with CCA-secure private-key encryption to obtain practical CCA-secure public-key encryption.

2. The Strong Diffie–Hellman assumption used in the lecture

2.1. A warning about terminology

The lecture first warns that the name Strong Diffie–Hellman is overloaded. Searching for this name online may lead to a different assumption, often a variant in which the adversary receives several powers related to a secret exponent and must compute an element involving an inverse of that exponent. That is not the formulation used here.

In this course, Strong Diffie–Hellman means a computational Diffie–Hellman problem augmented with a static-key DDH oracle. The slide therefore describes it as

CDH with a static-key DDH oracle.

The assumption is stronger than ordinary CDH in the sense that the adversary is given extra help. The claim that the resulting problem remains hard is therefore a stronger hardness assumption.

2.2. Comparison with ordinary CDH

Let \(\mathbb G\) be a cyclic group of prime order \(p\), generated by \(g\). In the ordinary Computational Diffie–Hellman problem, the adversary receives

\[ g,\qquad g^x,\qquad g^y \]

and must compute

\[ g^{xy}. \]

The course’s Strong Diffie–Hellman experiment gives the adversary exactly the same computational goal, but it additionally gives access to an oracle tied to the fixed, or static, exponent \(x\).

The adversary may submit two group elements

\[ g^r,\qquad g^s. \]

The oracle replies with the bit

\begin{equation*} \mathcal O_x(g^r,g^s) = \begin{cases} 1,& s=xr \pmod p,\\ 0,& \text{otherwise}. \end{cases} \end{equation*}

Equivalently, it tests whether

\[ g^s=(g^r)^x. \]

Thus, if the adversary submits \(g^y\) together with some candidate \(Z\), the oracle says whether

\[ Z=(g^y)^x=g^{xy}. \]

The oracle does not compute \(g^{xy}\) for the adversary. It only recognizes a correct candidate. It is also not a completely general DDH oracle accepting an arbitrary triple. One exponent, \(x\), is fixed throughout the experiment.

Formally, the adversary succeeds when it outputs \(Z=g^{xy}\). The SDH assumption used in the lecture states that, for every PPT adversary \(\mathcal A\),

\[ \Pr[\mathsf{SDH}_{\mathcal A}(\lambda)=1] \]

is negligible in \(\lambda\), even though \(\mathcal A\) has access to the static-key oracle \(\mathcal O_x\).

3. Reminder: Hashed ElGamal encryption

Let the secret key be \(x\in\mathbb Z_p\), and let the public key contain

\[ h=g^x. \]

For a message \(m\in\{0,1\}^{\kappa}\), encryption samples \(y\xleftarrow{\$}\mathbb Z_p\) and outputs

\[ c=(c_1,c_2) =\left(g^y,\ H(h^y)\oplus m\right). \]

Since

\[ h^y=(g^x)^y=g^{xy}, \]

decryption computes

\[ m=c_2\oplus H(c_1^x). \]

Indeed,

\[ c_1^x=(g^y)^x=g^{xy}=h^y. \]

The hash function is modeled as a random oracle in this proof.

4. The theorem: Hashed ElGamal is IND-CCA1 secure

The theorem stated in the lecture is:

Assume that the course’s SDH assumption holds in \(\mathbb G\). Then Hashed ElGamal is IND-CCA1 secure in the random-oracle model.

CCA1 means that the adversary may ask decryption queries before receiving the challenge ciphertext, but it receives no decryption oracle after the challenge. This is weaker than full CCA security, normally called CCA2.

The proof is a reduction from breaking the SDH assumption. It also illustrates three reusable proof techniques:

  1. extracting a hidden computational value from a random-oracle query;
  2. programming and lazily sampling a random oracle;
  3. simulating a decryption oracle without knowing the secret key.

4.1. The crucial observation: a successful adversary must query the hidden key

The challenge ciphertext has the form

\[ c^*=\left(g^y,\ H(h^y)\oplus m_b\right) =\left(g^y,\ H(g^{xy})\oplus m_b\right). \]

Let

\[ Q:=\text{the event that the adversary queries the random oracle at }g^{xy}. \]

Suppose the adversary never makes this query. Then, from its point of view, \(H(g^{xy})\) is a uniformly random, previously unseen string. Consequently,

\[ H(g^{xy})\oplus m_b \]

is a one-time pad encryption of \(m_b\). The challenge bit is then perfectly hidden, so the adversary can do no better than random guessing:

\[ \Pr[b'=b\mid \neg Q]=\frac12. \]

Equivalently, its distinguishing advantage conditioned on \(\neg Q\) is zero. Therefore, every adversary having non-negligible advantage must, with relevant probability, query the random oracle on exactly the hidden Diffie–Hellman element \(g^{xy}\).

This is the extraction point of the proof. The reduction cannot compute \(g^{xy}\) by itself, but it can wait for the Hashed ElGamal adversary to hand that element to the random oracle.

4.2. Constructing the SDH reduction

Assume there is an IND-CCA1 adversary \(\mathcal A\) against Hashed ElGamal. The reduction \(\mathcal R\) receives an SDH instance

\[ g,\qquad g^x,\qquad g^y, \]

and has access to the static-key oracle \(\mathcal O_x\). Its task is to output \(g^{xy}\).

The reduction sets

\[ pk=(g,g^x) \]

and gives this public key to \(\mathcal A\). The public key is distributed exactly as in the real Hashed ElGamal experiment.

Because this is a CCA1 game, \(\mathcal A\) may now submit polynomially many prechallenge decryption queries. Their simulation is discussed separately below.

Eventually, \(\mathcal A\) submits two equal-length challenge messages \(m_0,m_1\). The reduction samples

\[ b\xleftarrow{\$}\{0,1\} \]

and must create an encryption of \(m_b\). A genuine ciphertext would be

\[ \left(g^y,\ H(g^{xy})\oplus m_b\right), \]

but computing \(g^{xy}\) is precisely the reduction’s unsolved SDH task.

Instead, \(\mathcal R\) samples an independent random string

\[ R_H\xleftarrow{\$}\{0,1\}^{\kappa} \]

and constructs

\[ c^*=\left(g^y,\ R_H\oplus m_b\right). \]

It conceptually programs the random oracle so that

\[ H(g^{xy})=R_H. \]

The unusual feature is that the reduction knows the output \(R_H\) before it knows the corresponding input \(g^{xy}\). The static-key DDH oracle makes this possible: whenever the adversary later submits a random-oracle query \(Z\), the reduction tests

\[ \mathcal O_x(g^y,Z). \]

If the answer is one, then

\[ Z=(g^y)^x=g^{xy}. \]

The reduction has both recognized the preprogrammed random-oracle input and obtained the solution to its SDH challenge. It can output \(Z\) immediately. If the simulation is conceptually continued, the answer to this random-oracle query must be \(R_H\), ensuring consistency with the challenge ciphertext.

4.3. The reduction as a random-oracle intermediary

A student described the reduction as a kind of man in the middle between the adversary and the random oracle. The lecturer confirmed this interpretation. The adversary does not contact an independent random oracle behind the reduction’s back. In the reduction proof, \(\mathcal R\) implements the random-oracle interface seen by \(\mathcal A\).

This does not mean that the reduction may answer arbitrarily. Its answers must have the same distribution and consistency properties as a true random function. A convenient mental model is a table of input–output pairs. Previously unseen inputs receive fresh random outputs, and repeated inputs receive the same stored outputs.

Here, the reduction additionally creates a special output \(R_H\) in advance. The corresponding input is initially unknown. The SDH oracle allows the reduction to recognize that input when the adversary eventually supplies it. Since \(R_H\) itself is uniformly random, this programming is consistent with the distribution of a random oracle.

4.4. Simulating prechallenge decryption queries

The more delicate part is answering decryption queries without knowing the secret exponent \(x\). Suppose the adversary submits

\[ c=(A,B). \]

Real decryption would return

\[ B\oplus H(A^x). \]

The reduction maintains a random-oracle database containing known pairs

\[ H(D_i)=E_i. \]

For every existing input \(D_i\), it asks the static-key oracle whether

\[ D_i=A^x. \]

This is exactly the query

\[ \mathcal O_x(A,D_i). \]

If one entry satisfies this test, then \(E_i=H(A^x)\), and the reduction can return the correct plaintext

\[ m=B\oplus E_i. \]

The XOR works because, for a valid Hashed ElGamal ciphertext,

\[ B=H(A^x)\oplus m, \]

and therefore

\[ B\oplus H(A^x)=m. \]

There is also a case in which the adversary asks to decrypt \((A,B)\) before it has ever queried the random oracle at \(A^x\). A real random oracle still has some uniformly random value at that point, although nobody has requested it yet. The simulator therefore samples a fresh random string \(E\), returns

\[ B\oplus E, \]

and records a pending assignment saying, in effect,

\[ H(A^x)=E. \]

It does not know the group element \(A^x\), so the table entry initially cannot be indexed by that value. If a later random-oracle query \(D\) arrives, the reduction can test \(\mathcal O_x(A,D)\). If the test succeeds, it binds the pending output to the now-revealed input and returns the stored value \(E\). Thus the decryption answer and all later random-oracle answers remain consistent.

This simulation handles both honestly generated ciphertexts, for which the adversary normally queried the relevant hash input, and arbitrary bitstrings submitted as ciphertexts.

4.5. Why the simulation remains polynomial-time

A student asked whether searching through the random-oracle table would require examining exponentially many possible values. It does not. The reduction stores only the queries actually made by the adversary, together with a polynomial number of pending values created while answering decryption queries. Since a PPT adversary can make only polynomially many queries, the reduction searches only a polynomial-size table.

The implementation could also cache the results of static-key oracle tests to avoid repeating work, but this is an efficiency detail rather than a change to the argument.

4.6. Finishing the proof for all adversaries

The proof first focuses on adversaries that do query \(H(g^{xy})\). For such an adversary, the reduction recognizes that query and outputs the queried group element as its SDH solution. Hence a non-negligible advantage in this case would contradict the SDH assumption.

However, a security proof must cover arbitrary adversaries, including those that never make the critical query. Let

\[ p=\Pr[Q]. \]

Using the law of total probability,

\begin{equation*} \begin{aligned} \Pr[\mathcal A\text{ wins}] &=\Pr[\mathcal A\text{ wins}\mid Q]\Pr[Q]\\ &\quad+ \Pr[\mathcal A\text{ wins}\mid\neg Q]\Pr[\neg Q]. \end{aligned} \end{equation*}

Conditioned on \(Q\), any advantage beyond one half can be converted into an SDH attack, so under the SDH assumption this probability is at most \(\frac12+\nu(\lambda)\) for a negligible function \(\nu\). Conditioned on \(\neg Q\), the challenge is one-time-pad protected, so the probability is exactly \(\frac12\). Consequently,

\begin{equation*} \begin{aligned} \Pr[\mathcal A\text{ wins}] &\leq \left(\frac12+\nu(\lambda)\right)p +\frac12(1-p)\\ &=\frac12+\nu(\lambda)p\\ &\leq\frac12+\nu(\lambda). \end{aligned} \end{equation*}

Because a probability \(p\leq 1\) cannot turn a negligible function into a non-negligible one, the overall advantage is negligible. This proves IND-CCA1 security.

4.7. General lesson about reductions and simulation

The lecturer emphasizes that a reduction must simulate the real security experiment honestly from the adversary’s point of view. The adversary should not be able to notice that it is interacting with a reduction rather than the actual challenger and random oracle.

The lecture informally compares this requirement to a Turing test: the simulation must be behaviorally indistinguishable from the real system. If the adversary could detect that it was being used inside a reduction, it could simply stop behaving as expected and output nonsense. The lecturer also mentioned the Volkswagen emissions-testing scandal as an analogy for a system that behaves differently when it detects that it is being tested.

The proof techniques to remember are therefore:

  • identify a query or event that every successful adversary must cause;
  • use that event to extract the solution to the underlying hard problem;
  • program the random oracle while preserving the distribution of a random function;
  • maintain a consistent query table;
  • simulate every oracle given in the real security game;
  • finally remove any temporary restriction on the adversary by conditioning on the relevant event.

5. Why move to hybrid encryption?

After completing the CCA1 proof, the lecture turns to the practical problem of obtaining efficient, fully CCA-secure public-key encryption.

The first observation is that public-key operations are expensive. ElGamal, RSA, and similar constructions use modular exponentiation or analogous group operations, which are much slower than modern symmetric primitives. The slide gives illustrative rates of approximately

\[ 700\ \text{MB/s per core for AES} \]

and

\[ 1.5\ \text{MB/s per core for RSA encryption}. \]

The lecturer notes that these figures are hardware-dependent and may be old, especially because modern processors often contain dedicated AES instructions. The RSA figure refers to a secure practical RSA encryption construction rather than insecure textbook RSA. The exact numbers are not the point; the important fact is the large performance gap between symmetric and public-key encryption.

Encrypting a very short value directly with public-key encryption may be acceptable. Encrypting a large file, such as a 4K video, entirely with public-key operations is unnecessarily slow and may also produce poor ciphertext expansion.

Hybrid encryption combines the main advantage of public-key encryption—the sender only needs the recipient’s public key—with the speed of private-key encryption.

6. Basic hybrid encryption using PKE and SKE

Assume Alice wants to send a large message \(m\) to Bob and knows Bob’s public key \(pk\), but they do not already share a symmetric key.

Alice performs the following operations:

  1. Sample a fresh symmetric key

    \[ K\xleftarrow{\$}\{0,1\}^{\lambda}. \]

  2. Encrypt the short key using Bob’s public-key encryption scheme:

    \[ c_1=\mathsf{PKE.Enc}(pk,K). \]

  3. Encrypt the large payload using the fast private-key scheme:

    \[ c_2=\mathsf{SKE.Enc}(K,m). \]

  4. Send the combined ciphertext

    \[ c=(c_1,c_2). \]

Bob first recovers the symmetric key,

\[ K=\mathsf{PKE.Dec}(sk,c_1), \]

and then decrypts the payload,

\[ m=\mathsf{SKE.Dec}(K,c_2). \]

A symmetric key may be only 256 bits long, as with AES-256. Therefore the expensive public-key algorithm is used only once on a very short input, while the large payload is processed by the efficient symmetric scheme. The lecture uses Hashed ElGamal with a 256-bit hash output as an example that can transport an AES-256 key.

6.1. Whether to use a fresh public-key encapsulation for every message

A student asked whether this process must be repeated for every individual message or may instead be used per session. The lecturer’s answer is that this is partly a protocol and key-management policy decision.

One may generate a fresh key for each message. Alternatively, after the parties establish one shared key, they may derive subsequent keys from it and rotate those derived keys according to the protocol’s policy. The formal presentation in this lecture focuses on the initial situation in which Alice and Bob do not yet share any secret and Alice sends the first encrypted message using only Bob’s public key.

7. Efficiency advantages of hybrid encryption

7.1. Computational efficiency

Only the short key \(K\) is processed by the expensive public-key algorithm. The bulk payload is encrypted using a fast block-cipher-based or otherwise symmetric construction. The cost of public-key encryption is therefore nearly constant with respect to the payload length.

7.2. Ciphertext-size efficiency

For a typical symmetric encryption scheme, ignoring a comparatively small nonce, IV, or tag overhead, the payload ciphertext satisfies

\begin{equation*} |c_2|\approx |m|. \end{equation*}

The complete hybrid ciphertext has size

\begin{equation*} |c|=|c_1|+|c_2|\approx |c_1|+|m|. \end{equation*}

Therefore,

\[ \frac{|c|}{|m|} \approx \frac{|c_1|+|m|}{|m|} = \frac{|c_1|}{|m|}+1. \]

The public-key component \(c_1\) has fixed size for a fixed security parameter. For a very long message,

\[ \frac{|c_1|}{|m|}\longrightarrow 0, \]

and hence

\[ \frac{|c|}{|m|}\approx 1. \]

Thus, for something as large as a movie, the total ciphertext is essentially the same size as the symmetric encryption of the movie, plus one small public-key header. Hybrid encryption is efficient both in computation and in bandwidth.

8. From encrypted random keys to key encapsulation

In the basic construction, Alice chooses a random symmetric key and then uses public-key encryption only to transport that random value. Alice does not need to select a meaningful user message for this public-key operation. The cryptographic community therefore separates this narrower task into its own primitive: a Key Encapsulation Mechanism, or KEM.

The main conceptual change is:

  • ordinary public-key encryption takes a sender-chosen message as input;
  • encapsulation takes only the recipient’s public key and generates both a random key and a ciphertext encapsulating that key.

This can reduce ciphertext size and implementation cost because a KEM does not need to support arbitrary user-chosen plaintexts. The slide mentions that a factor of roughly two may be saved in constructions such as the Diffie–Hellman examples below.

An intuitive interpretation given in the lecture is:

Encapsulation is like encrypting a random message for Bob when Alice does not care what the message is, provided that Alice learns that random message and Bob can recover the same one.

The encapsulated key is not literally an input to the encapsulation algorithm; it is generated as part of the randomized computation. Nevertheless, the ciphertext enables the secret-key holder to reconstruct it.

9. Formal syntax and correctness of a KEM

A KEM consists of three PPT algorithms

\[ (\mathsf{KeyGen},\mathsf{Encaps},\mathsf{Decaps}). \]

9.1. Key generation

\[ (pk,sk)\leftarrow\mathsf{KeyGen}(1^\lambda). \]

This randomized algorithm generates a public key and a secret key, just as in public-key encryption.

9.2. Encapsulation

\[ (c,K)\leftarrow\mathsf{Encaps}(pk). \]

This is randomized. It takes no message input. Repeating it with the same public key normally gives a fresh ciphertext and a fresh key.

In the sender/receiver picture, Bob first generates \((pk,sk)\), and Alice runs \(\mathsf{Encaps}(pk)\). Alice learns both \(c\) and \(K\), but sends only \(c\) as the KEM part of the final ciphertext.

9.3. Decapsulation

\[ K'\leftarrow\mathsf{Decaps}(sk,c). \]

In the syntax presented in the lecture, decapsulation is deterministic. Bob uses his secret key and the encapsulation ciphertext to recover the key.

9.4. Correctness

Whenever

\[ (pk,sk)\leftarrow\mathsf{KeyGen}(1^\lambda) \]

and

\[ (c,K)\leftarrow\mathsf{Encaps}(pk), \]

correctness requires

\[ \Pr[\mathsf{Decaps}(sk,c)=K]=1. \]

Thus Alice’s encapsulated key and Bob’s decapsulated key are identical.

10. Why KEMs are the modern abstraction

The lecture notes that there is little practical motivation to encrypt a large payload directly with a public-key scheme when hybrid encryption gives the same public-key usability with much better performance. It is therefore natural to design only the public-key component needed for key transport.

The lecturer also connects KEMs to post-quantum cryptography. Public-key systems based on discrete logarithms and related classical assumptions are threatened by sufficiently capable quantum computers. In post-quantum standardization, it has often been easier and more useful to design a KEM than a general public-key encryption scheme supporting arbitrary plaintexts. Since real applications normally use hybrid encryption anyway, KEMs have become the standard public-key building block in this setting.

11. Examples of KEMs

11.1. A trivial KEM from public-key encryption

Any public-key encryption scheme can be used to construct a KEM:

\begin{equation*} \begin{aligned} \mathsf{Encaps}(pk):\quad &K\xleftarrow{\$}\{0,1\}^{\lambda},\\ &c\leftarrow\mathsf{PKE.Enc}(pk,K),\\ &\text{return }(c,K); \end{aligned} \end{equation*}

and

\[ \mathsf{Decaps}(sk,c): \quad K\leftarrow\mathsf{PKE.Dec}(sk,c). \]

This construction exactly formalizes the first hybrid-encryption picture. However, it does not exploit the possibility of designing a smaller primitive specifically for random keys. It still carries all the overhead of a full public-key encryption scheme.

The fact that PKE immediately gives a KEM suggests that general public-key encryption is a stronger primitive. The reverse direction is not obtained by the KEM alone; to encrypt arbitrary messages, it must be combined with a private-key encryption scheme.

11.2. Every two-message key exchange gives a KEM

The lecture observes that a two-message key-exchange protocol can be reinterpreted as a KEM.

In a two-message exchange, one participant sends an initial public value, the second participant sends a response, and both derive the same shared key. To view this as a KEM:

  • the first participant’s long-lived first-message state becomes the KEM key pair;
  • the second participant’s response becomes the encapsulation ciphertext;
  • the shared key computed by the second participant is the output of encapsulation;
  • the first participant’s computation of the same shared key is decapsulation.

Diffie–Hellman gives the central example.

11.3. Diffie–Hellman KEM

Key generation samples

\[ x\xleftarrow{\$}\mathbb Z_p, \]

sets

\[ h=g^x, \]

and outputs

\[ pk=(g,h),\qquad sk=x. \]

Encapsulation samples

\[ y\xleftarrow{\$}\mathbb Z_p \]

and computes

\[ c=g^y, \qquad K=h^y=g^{xy}. \]

It returns \((c,K)\).

Decapsulation computes

\[ K=c^x=(g^y)^x=g^{xy}. \]

This is precisely a Diffie–Hellman exchange with one side’s public value fixed as the KEM public key. The encapsulation ciphertext contains only the single group element \(g^y\), whereas ordinary ElGamal public-key encryption contains two group elements.

Its IND-CPA security follows directly from the eavesdropping-security experiment for Diffie–Hellman key exchange: the relevant experiments are essentially identical. The adversary must distinguish the real shared key from an independent random key given the public transcript.

11.4. Hashed ElGamal KEM

To obtain a bitstring key suitable for symmetric encryption, hash the Diffie–Hellman group element. Key generation is unchanged:

\[ h=g^x, \qquad pk=(g,h), \qquad sk=x. \]

Encapsulation samples \(y\xleftarrow{\$}\mathbb Z_p\) and computes

\[ c=g^y, \qquad K=H(h^y)=H(g^{xy}). \]

Decapsulation returns

\[ K=H(c^x). \]

Again, the KEM ciphertext is only one group element. In contrast, Hashed ElGamal public-key encryption also needs a second bitstring component \(H(h^y)\oplus m\).

12. Security definitions for KEMs

KEM security resembles the security of a key-exchange protocol more than the usual two-message IND experiment for encryption. The adversary is explicitly given either the real encapsulated key or a random key and must distinguish between them.

12.1. IND-CPA security for KEMs

The challenger generates

\[ (pk,sk)\leftarrow\mathsf{KeyGen}(1^\lambda) \]

and encapsulates

\[ (c^*,K)\leftarrow\mathsf{Encaps}(pk). \]

It samples

\[ b\xleftarrow{\$}\{0,1\}. \]

If \(b=0\), it sets

\[ K^*=K. \]

If \(b=1\), it sets

\[ K^*\xleftarrow{\$}\{0,1\}^{\lambda}. \]

The adversary receives

\[ pk,\qquad c^*,\qquad K^* \]

and outputs a guess \(b'\).

A KEM is IND-CPA secure if

\[ \left| \Pr[b'=b]-\frac12 \right| \]

is negligible for every PPT adversary.

This definition combines features of two earlier experiments. In public-key encryption, the adversary chooses two messages and receives an encryption of one of them. In key-exchange security, the adversary sees either the genuine session key or a random key. The KEM game gives the adversary the encapsulation ciphertext and asks whether the accompanying key is the real one encapsulated by that ciphertext or an unrelated random string.

This definition is exactly what hybrid encryption needs: if the encapsulated key is computationally indistinguishable from random, it can safely serve as the key of the symmetric payload-encryption scheme.

12.2. IND-CCA security for KEMs

The CCA game additionally gives the adversary a decapsulation oracle. It may submit ciphertexts \(c\) and receive

\[ \mathsf{Decaps}(sk,c). \]

In the full CCA, or CCA2, game, these queries are allowed both before and after the challenge, except that after receiving \(c^*\), the adversary may not ask to decapsulate that exact ciphertext:

\[ c\neq c^*. \]

No encapsulation oracle is needed. Encapsulation uses only the public key, so an adversary possessing \(pk\) can already run \(\mathsf{Encaps}(pk)\) by itself.

This is a strong security notion. Informally, the adversary may learn the keys corresponding to arbitrarily many chosen encapsulation ciphertexts and must still be unable to decide whether the key associated with the one forbidden challenge ciphertext is real or random.

A student asked again about the distinction between CCA1 and CCA2. CCA1 allows decapsulation or decryption queries only before the challenge. CCA2 also allows them afterward, excluding the exact challenge ciphertext. In the rest of this lecture, the unqualified name CCA security means the full CCA2 notion.

13. Hashed ElGamal KEM achieves full CCA security

Under the same SDH assumption and in the random-oracle model, the lecture states:

The Hashed ElGamal KEM is IND-CCA2 secure.

This is stronger than the result proved for Hashed ElGamal public-key encryption, which is only IND-CCA1 secure.

13.1. Why Hashed ElGamal PKE is not CCA2 secure

Write the Hashed ElGamal challenge ciphertext as

\[ c^*=(A,B) =\left(g^y,\ H(h^y)\oplus m_b\right). \]

After seeing it, a CCA2 adversary chooses a known nonzero mask \(\Delta\) and constructs

\[ c^{**}=(A,B\oplus\Delta). \]

This is not literally the forbidden ciphertext because

\[ c^{**}\neq c^*. \]

The adversary may therefore submit it to the postchallenge decryption oracle. The result is

\begin{equation*} \begin{aligned} \mathsf{Dec}(sk,c^{**}) &=(B\oplus\Delta)\oplus H(A^x)\\ &=m_b\oplus\Delta. \end{aligned} \end{equation*}

Since the adversary chose \(\Delta\), it recovers \(m_b\) and hence learns \(b\). This is a direct malleability attack on the second ciphertext component. Hashing the Diffie–Hellman value does not by itself authenticate the XOR-encrypted message.

13.2. Why the same attack does not apply to the KEM

For the Hashed ElGamal KEM, the challenge consists of

\[ c^*=g^y \]

and a separately supplied test key

\begin{equation*} K^*= \begin{cases} H(g^{xy}),&\text{real case},\\ U,&\text{random case}. \end{cases} \end{equation*}

There is no second ciphertext component of the form \(H(g^{xy})\oplus m_b\) that the adversary can modify while leaving the same Diffie–Hellman value underneath. Changing the single group-element ciphertext changes the value that will be decapsulated.

The proof again relies on the random-oracle query principle. Unless the adversary queries the oracle at \(g^{xy}\), the real key \(H(g^{xy})\) is indistinguishable from a fresh random string. Once the adversary does query that point, the SDH reduction recognizes the query with the static-key oracle and extracts \(g^{xy}\). The decryption-oracle simulation techniques from the CCA1 proof are reused, now in the KEM game.

Thus the KEM formulation removes the simple XOR malleability that prevents Hashed ElGamal PKE from achieving CCA2 security.

14. Public-key encryption from a KEM and private-key encryption

Let

\[ (\mathsf{KeyGen},\mathsf{Encaps},\mathsf{Decaps}) \]

be a KEM, and let

\[ (\mathsf{Enc},\mathsf{Dec}) \]

be a private-key encryption scheme whose keys are uniformly distributed in the key space expected by the scheme.

The combined public-key encryption scheme

\[ (\mathsf{KeyGen}',\mathsf{Enc}',\mathsf{Dec}') \]

is defined as follows.

14.1. Key generation

\[ \mathsf{KeyGen}'(1^\lambda): \quad (pk,sk)\leftarrow\mathsf{KeyGen}(1^\lambda). \]

The combined system simply uses the KEM key pair.

14.2. Encryption

To encrypt \(m\) under \(pk\), compute

\[ (c_1,K)\leftarrow\mathsf{Encaps}(pk), \]

then

\[ c_2\leftarrow\mathsf{Enc}(K,m), \]

and output

\[ c=(c_1,c_2). \]

The KEM component \(c_1\) transports a fresh key, while \(c_2\) contains the large encrypted payload.

14.3. Decryption

Given

\[ c=(c_1,c_2), \]

compute

\[ K\leftarrow\mathsf{Decaps}(sk,c_1) \]

and then

\[ m\leftarrow\mathsf{Dec}(K,c_2). \]

This is the same hybrid-encryption workflow introduced at the start of the second half of the lecture, except that a purpose-built KEM replaces the step of encrypting a sender-chosen random key with a general PKE scheme.

15. Security theorem for the hybrid construction

The final theorem states:

If the KEM is IND-CCA secure and the private-key encryption scheme is IND-CCA secure, then the combined construction is an IND-CCA-secure public-key encryption scheme.

Here, CCA means full CCA2 security.

The high-level reasoning is that KEM security lets the proof replace the encapsulated key by an independent random key without the adversary noticing. Once the payload is encrypted under an independent random key, the CCA security of the private-key encryption scheme hides the challenge message and prevents useful ciphertext modifications.

The lecture does not develop the complete reduction for this composition, but emphasizes the consequence: instead of designing a monolithic public-key scheme that directly encrypts arbitrarily large messages and is itself CCA2 secure, one combines two well-understood components:

  • a CCA2-secure KEM, such as the Hashed ElGamal KEM under the stated SDH and random-oracle assumptions;
  • a CCA-secure symmetric encryption construction, meaning an authenticated or otherwise chosen-ciphertext-secure mode rather than a bare block-cipher invocation.

The result is efficient CCA2-secure public-key encryption. This is the modern and practical approach: the public-key mechanism handles only a short fresh key, and the private-key mechanism handles the payload.

16. Final takeaways

The first half of the lecture completes the Hashed ElGamal CCA1 proof. A successful adversary must query the random oracle at the hidden Diffie–Hellman value \(g^{xy}\). The reduction uses the course’s SDH oracle to recognize that query, extracts the CDH solution, programs the random oracle, and simulates prechallenge decryption through a polynomial-size query table. Conditioning on whether the critical hash query occurs extends the argument to all adversaries.

The second half explains why large data should not be encrypted directly with public-key operations. Hybrid encryption uses public-key cryptography only to establish or transport a short symmetric key and uses efficient private-key encryption for the payload. Its ciphertext expansion approaches one for long messages.

A KEM captures exactly the public-key task needed by hybrid encryption: encapsulation outputs a fresh key together with a ciphertext, and decapsulation recovers the same key. Diffie–Hellman naturally yields a KEM, and Hashed ElGamal yields a bitstring-valued KEM with a one-group-element ciphertext.

Finally, Hashed ElGamal illustrates an important distinction. As a direct PKE scheme it is malleable and only CCA1 secure under the lecture’s theorem. As a KEM it can be fully CCA2 secure under the same SDH assumption in the random-oracle model. Combining such a CCA-secure KEM with CCA-secure symmetric encryption gives the efficient CCA-secure public-key encryption used in practice.

Author: Lowtroo

Created on: 2026-08-02 Sun 23:20

Powered by Emacs 29.3 (Org mode 9.6.15)