DOE Early Career Research Program, Office of Science, 2020–2026
Data-Driven Optimization under Uncertainty
Closing notes on a five-year award that took seven, and the question it did not settle.
The objection
I wrote the proposal around one objection. Almost every method for deciding under uncertainty begins by assuming you know the odds: estimate a distribution from past data, then find the plan that does best against it. My argument was that the distribution you estimated is not the one you will meet, and a plan tuned tightly to an estimate can fail badly when the estimate is only a little wrong.
So carry a whole set of plausible distributions instead of one, and choose the plan that survives the worst of them — across a sequence of decisions rather than a single shot. Then make it run on machines large enough to matter: thousands of pieces solved at once, on GPUs, tolerating pieces that arrive late or only half-solved.
The proposal named two places to try it. The first was the electric grid. The second, two pages near the back, was training models across many small computers that could not share their data.
This sensitivity, known as the optimizer's curse, is due to overfitting effects, as also seen in machine learning. Project narrative, April 2019
I did not think much about that sentence when I wrote it.
What I did not expect
The second application took over.
It began with a differentially private federated learning method built on inexact ADMM — an algorithm from the award, aimed at the problem the award had treated as an afterthought. It transferred almost without modification, and once it did, the centre of the work moved with it.
I could call that a project that drifted. I would rather say what actually happened, which is that the mathematics turned out to be about something larger than the frame I had put around it.
Why it was not a detour
The set of plausible distributions was never really about probability. It was about the gap between the data I have and the world I work in.
Federated learning is that gap made physical. Every hospital, every X-ray facility, every utility holds one slice of the data. No slice is the truth. The model has to work for all of them. And the same machinery applies: a decomposition method splits one large problem into pieces solved separately and reconciles them through a coordinating step; federated learning trains copies of a model at sites that will not hand over their data and reconciles them the same way.
What the award leaves behind
The mathematics held on its own terms. The theory of splitting distributionally robust integer problems and putting the answers back together was given a semi-plenary at the International Conference on Stochastic Programming — the community the proposal had been written for, saying the work had landed. The promise to use GPUs was kept, though not in the shape I imagined: instead of accelerating one step inside a decomposition method, we moved whole power flow calculations onto them.
APPFL, the open software for privacy-preserving federated learning, came out of the second application and is still maintained. It is the thing people actually picked up.
Four postdoctoral researchers came through the project — Geunyeong Byeon, Minseok Ryu, Hideaki Nakao and Charikleia Iakovidou — along with about twenty visiting graduate students. Geunyeong and Minseok are both now on the faculty at Arizona State, and Minseok is a co-investigator on what followed.
The award was funded for five years and ran for seven. Careful spending left carry-over, and two no-cost extension years followed. They did not behave like a coda: they produced the journal papers the earlier years had set up, and they overlapped with the beginning of everything that came after. For two years the award was running at the same time as the work that replaced it. The question was never handed over at a boundary. It changed names while the original grant was still paying for it.
What it handed off to
In September 2024, partway through the extension, I spoke to DOE's Advanced Scientific Computing Advisory Committee and called the talk Building on Success: Advancing Privacy-Preserving Federated Learning with Distributed Optimization. I was making the case for the successor while the original was still running.
That successor is Privacy-Preserving Federated Learning for Science, part of DOE's AI for Science program: building foundation models across DOE facilities without moving any facility's data. Around it now sit foundation models for the grid, two Genesis Mission projects, a CESER project on privacy in grid operations, and a CMEI project on interconnecting large loads.
There is an inversion in this I did not see coming. The proposal put distributed learning at the edge of the network: many small computers, too little bandwidth to send their data anywhere. The work now runs at the other end of the scale — federated learning across DOE supercomputers, training one model on several leadership machines at once. It also cashes the part of the proposal I thought was about something else. I had promised algorithms that keep working when some pieces arrive late or only half-finished; when the pieces are jobs waiting in the batch queues of different facilities, that tolerance stops being a refinement and becomes the whole problem.
The hard part has not moved. A model trained on the grids we have will be asked about a grid it has never seen. The data it learned from is not the world it will be used in. In 2019 I called that distributional ambiguity. Now I call it generalization.
It ended this year. An Early Career award buys five years to be wrong about the details — seven, in my case. My error was one of filing. I listed distributed learning across sites as an application, a place to try the method out. It was not an application. It was a second home for the question, and the question is the part I got right.
What the question has drawn
Counting the Early Career award and the work that grew from it: privacy-preserving federated learning from PALISADE through Federated Learning for Science, the grid foundation-model line, two Genesis Mission projects, the CESER privacy project and the CMEI large-load project. Figures are total project value, including work led elsewhere on projects where I am a co-investigator.
- Argonne's announcement of the award, August 2019
- APPFL — privacy-preserving federated learning
- LUMINA-2M — grid foundation model on Hugging Face
- Publications on Google Scholar
- Back to the main page