Qubes Ghost: amnesic session, portable qubes on encrypted removable media, saved forward only

I ran the test I promised. Rolling back one member of a group, on SimpleX.

Setup: three profiles, alice, bob and carol, one group. alice writes, bob reads, carol is the third member and she is there as a control. I copy bob’s state directory, that is snapshot one. alice writes more, bob reads them fine. Second snapshot. Then I put the first snapshot back over bob’s state, start him again, and alice keeps writing.

Result. carol received both of the new messages. bob received none. Before the rollback he had read everything, the messages from before the snapshot and the ones after it.

So the conclusion is simple. Snapshot and restore do not get along with crypto that only moves forward. A qube with a messenger like that can be saved forward, but it must never be brought back from an older archive.

The method deserves a separate note, it may be more useful than the result itself. My first two runs gave me a false confirmation. The script printed “the rolled back member could not read”, but in reality nobody was reading anything, because the group was not delivering at all. What caught it was the control member, the one I never rolled back. If I had used only two profiles I would be posting nonsense right now. The real cause turned out to be the syntax of the send command, I wrote it from memory and got it wrong, and the client stayed silent instead of returning an error. So I added two gates to the script: do not go further until it is proven that the contacts are established and that messages actually arrive before the snapshot.

I only tested SimpleX. I am not claiming anything about Matrix, I will set that up separately and report.

One more thing. I read the live mode thread, dom0 in RAM, and the neighbouring ones about overlays and the encrypted pool. I am taking part of it. An ephemeral dom0 removes a whole class of problems that I was fixing with patches: leftover logs, panel state, metadata. There it is gone by construction, while I was cleaning it by hand and still not completely.

But I am not taking all of it. Full amnesia breaks exactly what this test showed. If the messenger state disappears on shutdown then every boot is a new identity and conversations do not survive. So mine will be mixed. The session is ephemeral, and the working qubes come from removable encrypted media and are saved back to it, forward only.

1 Like

@newqube thank you for pointing me at that thread. I spent today actually reading it, together with the overlay one and the encrypted pool one next to it.

I will say it plainly. On the part where our approaches overlap, yours is much better than mine. I was cleaning dom0 by hand, log files, panel state, metadata, one path at a time, and I kept finding new ones. With a read only root and an ephemeral overlay there is nothing left to clean. That is not a better patch, it removes the problem.

So I changed the concept, not a detail of it. The separate RAM pool I had built goes away and the session itself becomes ephemeral. Several things I had open as bugs stop being bugs, including the one where my scrubbing only unlinked files instead of erasing them. A fix I spent hours on this morning is now simply unnecessary.

I am not closing the project though, because there is a layer your modes do not cover, and I only understood how much it matters after running the test I posted above. Full amnesia kills any messenger whose crypto only moves forward. Every boot becomes a new identity and conversations do not survive. That is not a corner case for me, it is the main thing I use these machines for.

So the work continues on the other layer. Workload qubes that live on encrypted removable media, get loaded into the ephemeral session, stay physically detached while I work, and are saved back forward only. Deniability about whether that media carries anything at all is also still mine to build.

Short version: your part is the system, my part is what the system carries. I would rather build on top of yours than keep patching my own worse version of it.

@James369 you asked the same thing earlier in this thread, whether the clean way is to make the whole system ephemeral, dom0 and all VMs. You were right and I talked past it at the time. Sorry about that.

1 Like

Qubes does write - logs, qubesdb, menu files…

2 Likes

still could put more effort in compacting LLM responses to something more meaningful :slight_smile:

1 Like

Credit correction. The amnesic mode I build on is linuxuser1’s live mode thread, Qubes OS live mode. dom0 in RAM. Non-persistent Boot. RAM-Wipe. Protection against forensics. Tails mode. Hardening dom0. Root read‑only. Paranoid Security. Ephemeral Encryption, not newqube. He pointed me there, the work is linuxuser1’s, I should have named the thread.

As a separate project mine does not hold up any more. The overlapping part is done better there and my version would just be a worse copy. So I stop it as its own thing. Bugs I hit in the overlay scripts and anything I can add go to that thread, not to mine. I have two fixes from running the modes on real hardware, they go there.

The one part where my approach still has a point is the air gap. The working material should not sit on the machine at all. It comes from removable media, the media is physically out while the session runs, then it goes back. An amnesic session alone does not give that.

So I do it inside the same setup instead of a separate one. The system goes on the removable media too and the key material is not on the machine. I will describe it after I run it, not before.

2 Likes

That sounds really interesting, I look forward to it. It may solve some important issues.

1 Like

We are now chatting with language models in the forum, eh? Interesting.

3 Likes

Correction to something I said here. I wrote that the separate RAM pool goes away. That was too broad and it is wrong.

It was only redundant as a way to make dom0 amnesic, which linuxuser1’s live mode does properly. As the place the workload actually runs it is still the thing I am building, and it is back. The cycle is: restore the qubes from the encrypted store into a pool in RAM, close the store so it is not in the system at all, work, open it again only to save. During the session the disk sees nothing.

Ran it today on the laptop, not in a VM. Wrote a marker in a qube, closed the store, the qube kept running and I wrote more into it with the store gone, saved, tore the whole RAM pool down, brought it back up, restored. All of it was there.

The script that does the cycle is here, and it is the one I actually ran, not a cleaned up version: qubes-ghost/scripts/ghost at main · Qubes-Ghost/qubes-ghost · GitHub

Two things that cost me time, in case they save someone else’s. The thin pool in RAM has to be deactivated and reactivated before a restore or LVM refuses it, “prohibited while rpool_tmeta is active”. And qvm-backup-restore has no option for the target pool, so you have to switch the default pool for the duration and put it back.

One constraint this puts on the design. A qube that works this way has to be standalone, or its template must be in RAM too. If the template stays in the store you cannot close the store, the qube loses its root. Anything too big for RAM runs the other way with the store open, which for me is fine, the big things hold public data.

1 Like

Matrix, as promised. It does not break like SimpleX.

Same test. Three accounts in an encrypted room, I roll one back to an older copy of its state, the third one untouched as a control. My own server, matrix-nio as the client.

The rolled back one read the message sent after the rollback. So did the control. Nothing arrived undecrypted. Group session is why: if your ratchet state goes back you can wind it forward again and read later messages of that session, you just cannot read earlier ones. SimpleX breaks because the sender never gets a new key from the rolled back side.

Added a fourth member to force a new session, same result. I do not know whether it actually rotated in my client or the missing key was just re-requested. Do not take that part as tested.

First run of this was wrong and I nearly posted it. Every restart was a fresh login, so a new device with new keys, and the rolled back one was failing before I rolled it back. The control caught it. If you run this, restore the same device from the store instead of logging in, and keep the credentials inside the snapshot.

Nothing changes for the rule. Reading survives, forward secrecy does not: after a rollback the other side never sees a new key, so whoever took the state keeps reading. Save forward only.

1 Like

A correction to what I said about Matrix.

I wrote that I added a fourth member to force a new session and got the same result. That was wrong. The run itself was no good: the test accounts still carried devices from earlier runs with no one time keys left, so sharing the key for a new session broke on its own, before any rollback. I marked that part untested at the time, but I should not have written “same result” either.

Redone. Fresh accounts for every run, events counted only in the test room, and the control gate now stops the test instead of printing a line into the log. And I read the session id now, instead of inferring a rotation from the fact that the membership changed.

The session does rotate. One id before the fourth member joined, another after, and the message sent after the rollback used the new one.

The rolled back account did not read it. One undecryptable event. The untouched control read it fine. A run without the rotation confirms the other half: there the same rollback costs nothing.

No key re-request went out on its own. That client has a single device and no key backup, though, and a re-request is answered by the user’s other devices or by the backup, so there was nothing to ask. A real client with key backup enabled may well heal where this one did not. I have not tested that and I am not claiming it.

So, replacing what I said before: Matrix survives a rollback up to the next change of session, and loses reading after that. It does not soften the forward only rule, it supports it.

1 Like

Hi, just another voice in the chorus-- your LLM-wrapped communication and the sequential posting of corrections does not engender confidence in your work. It causes the reader (well, this reader) to wonder if you possess the experience and expertise to be working on a project of this kind, where a bug could have major consequences to to the user. As why else would they want to be using an ‘amnesic’ system, if not for safety from some true threat?

In general, LLM-generated code is only good code if it’s being guided, or reviewed, by a subject matter expert.

I’m not sure if my post is meant to be a warning to a potential user, or advice to you the contributor, perhaps both.

1 Like