Skip to main content
Back to articles
2026 / 10
| 7 min read

Catching up on messages over a radio mesh

Thinking through how a MeshCore message store could fill in a missed conversation, and how devices would find it without using up the shared airtime.

meshcore radio distributed systems design notes
On this page

I got home and opened my phone to see what people had been talking about on the MeshCore public channel that day. I could see enough to be curious, but it felt like I had only heard half of a phone call. There were pieces missing. Now that I wanted to ask about them I had poor coverage, nobody to ask, and no history to go back through.

When I’m out around Edmonton I have coverage, but I’m usually out doing something. Home is where I have time to read, or where an idea turns up that I’d like to share with someone, and it is also where the connection is less reliable. Putting up more repeaters is interesting work and helps with coverage, but I started wondering about what those little computers on the masts could remember between the times we hear from one another. If a message passed through while I was away, could some part of the network keep it until I was able to receive it?

We’re fairly used to this working with text messages. We don’t have to arrange to be connected at exactly the same moment as the person sending one, because there is storage along the way. I’d like the mesh to be able to hold a message for someone who is out of reach. Keeping a few messages in memory seems quite manageable; finding out who has them, and getting them to the person who missed them, needs more thought when all of us are sharing a slow radio channel.

Suppose a repeater kept the recent messages it heard on a public channel. When my device came back into range, it could ask for the part of the conversation it hadn’t received. The repeater would need some way to recognize what I already had, so that it could send the missing messages without sending the whole conversation again. Here is a small example, using letters as message identities:

PlaceMessages it has
My deviceA, B, D
Nearby storeA, B, C, D
Useful replyC

Saying “I have D” wouldn’t describe that gap. It might mean that I had received everything through D, or that I had missed something earlier and happened to hear D. We would need a way to distinguish those cases. The letters aren’t a proposed packet format or a global sequence number; they just let us see what information the two sides need before deciding how to encode it.

Even with one device and one store, there is still the question of how much history to keep and what to do if the connection drops partway through catching up. A retry should be able to pick up from what arrived, which brings us back to being able to identify individual messages.

With one device and one store, at least there is only one place to ask. The next question is what two stores look like, or four, when the device can hear more than one of them. Those stores might have heard the same conversation, or different parts of it, so there is more to work out than simply adding another copy of the first store.

Finding messages without asking everyone

If every store answers the same request, several repeaters could all try to send me the same missing conversation. Having more copies would have made it easier to find the messages, but I would be paying for those copies in airtime even though one answer would have done. If the request itself is flooded across the mesh, there is also the traffic spent forwarding it before anybody gets as far as answering.

Imagine 90 repeaters in a small area, with a thousand users on a network moving perhaps two to five kilobits per second. These are numbers to think with, rather than measurements of the Edmonton network. A single device checking for new messages doesn’t sound like much, but we have to account for how often it asks, how many times the question is repeated, and how many answers come back. Most of the devices might already have everything on the channels they are following and we would still have spent the time finding that out.

We can put some numbers around just the asking. Suppose each of those thousand devices sends a 96-byte check-in, and suppose, for this calculation, that all the transmissions compete for the same channel capacity. With one copy of each check-in, the amount of data is:

1,000 devices × 96 bytes × 8 bits/byte = 768,000 bits
768,000 bits ÷ 2,000 bits/second = 384 seconds

That would take 6.4 minutes of continuous transmission at 2 kbit/s. Spread it over an hour and it uses about 11% of the capacity in this simplified model; ask every five minutes and the check-ins alone require more capacity than is available.

Check-in intervalAt 2 kbit/sAt 5 kbit/s
Every 60 minutes10.7%4.3%
Every 15 minutes42.7%17.1%
Every 5 minutes128.0%51.2%

These are calculated fractions of the shared bit capacity, not LoRa airtime measurements. The 96 bytes are an assumed size, and this leaves out radio framing, preambles, contention, retries, replies and the messages we actually wanted. Transmitting three copies of a check-in through that same contention domain would triple its contribution. A real mesh has geography and spatial reuse, so multiplying everything by all 90 repeaters would be a different, and rather poor, model. Even this little calculation is enough to show why the interval and the number of copies matter.

This is where I’ve been thinking about lightweight check-ins, so devices and stores know where to try to rendezvous. For the public channel I’m describing, a store wouldn’t know which part of the conversation I was missing until my device told it. The check-in could identify the channel and give some indication of what I had already received; a store could then work out whether it had anything useful to offer. There would still need to be a way to choose who answers when several stores have those messages. We need to be able to find one another without having to yell at everything and have everything yell back for every state change.

A possible catch-up exchangeA returning device describes its channel and received history. A store is selected through an as-yet undecided mechanism. Missing messages are sent in a bounded transfer, and the device records progress for a retry.

Yes, when capacity allows

No

Device returns to coverage

Describe channel and received history

Choose a reachable store

Transfer a bounded part of the gap

Record what arrived

More wanted?

Stop

The box labelled “choose a reachable store” is doing quite a lot of work in that drawing. I haven’t picked a discovery or election protocol for it. It identifies the part that needs resolving before several stores can help without all answering at once.

There is a cost to keeping that knowledge current, too. A device moves, a path stops working, or the store it was talking to is no longer reachable. Registering every movement everywhere could become as chatty as repeatedly asking for messages. How much do the nearby devices need to know, and how much needs to travel further through the mesh? I haven’t settled that, and it affects how useful these check-ins could be. Something that works well while a device stays in one place might create a lot more traffic as people come and go.

How much conversation to keep

Once there is somewhere to leave messages, it is tempting to keep everything, but that also leaves us deciding how much old traffic to send when somebody returns. I wanted enough history to understand what people were discussing. That doesn’t necessarily mean I want five days of channel traffic arriving over the radio when I’ve finally found a usable connection.

A store could retain recent messages and let the older ones expire. The awkward part is choosing what recent means for the conversation and for the capacity available to carry it. A device that has been away for an hour might be missing only a few messages; another that has been away for days might have a much larger gap. I’d like to be able to bound how much work either request creates, while still giving the person some indication of what history is available. Otherwise a growing backlog could occupy the channel just as people are trying to have a new conversation.