Somebody’s waiting for a bus and didn’t know it was gonna be late, and they got late, to a job interview. Like, it just totally changed the trajectory of their, like, life. Stop Requested.
This is Stop Requested. by ETA. I’m Christian. And I’m Levi. These are real conversations with the innovators, operators, and advocates driving improvements in public transportation.
Today, we’re continuing our AI roundtable series with ETA. CEO John Maglio and Derek VanGenneb, a quantitative researcher at Volridge Investment Management. Transit agencies collect enormous amounts of data about what has already happened, but what does it take to use that data to predict what happens next? We’ talk about how predictive models are built and tested, why data quality matters so much, what happens when conditions change, and where machine learning could help transit agencies anticipate everything from ridership and maintenance needs to service performance.
Here’s our conversation with Derek VanGenneb, and John Maglio. Welcome back to Stop Requested. We’re back with another great episode, continuing our AI roundtable series, and I wanna start by introducing our panel this morning, or, you know, everybody can introduce yourselves so our listeners know who’s part of this great conversation. Your co-host, Christian Londono with ETA.
And I’m your co-host, Levi McCollum, also with ETA. John Maglio, CEO, ETA. My name is Derek VanGenneb. I’m a quantitative researcher at Volridge, which is a, which is a quant finance hedge fund. Derek, thank you for joining us, and, and I wanna start, you know, with you, asking you this question.
So you work, your work comes from a different environment than public transportation. All of us, we’re kinda like in the public transit world. But, but before we jump into transit specifically, how do you think about using data and machine learning to make predictions about what’s going to happen next?
It’s a good question. First, thank you guys, thank you guys for having me. It’s, it’s, it’s good to be here. So I think, you know, in terms of like using data to predict the future, it all starts with having good data.
Like, you’re– let’s say you built a machine learning model, it only knows what, what you give it. And if you don’t have good data to train these models and to use them in the future, then the models aren’t that useful in, in predicting things in reality.
So, you know, first I try to think really deeply about what information is contained in the data that I have, and maybe also about like what data don’t I have and what data do
I think would be really useful for whatever I’m trying to predict. And then, you know, once you have what you believe to be good data, the sort of key test in figuring out whether or not you can predict what’s going to happen in the future is to be really careful with how you train these machine learning models and how you test them. So you might train some machine learning model on some chunk of data, and then test it on some data that you or the model hasn’t seen yet. And it’s like super important to do this in a, in a careful way where, where you’re not, you’re not cheating, you’re not introducing any like leakage or f- you know, forward-looking data into the model and, and kind of, you know, cheating it in a certain way.
So, you know, in, in finance, we– it’s kind of a, an extreme example. We have lots of noisy, changing data, you know. These distributions of, uh, different variables are constantly changing and, uh, we have to sort of constantly adapt. And so, um, things like, you know, making sure you have good data, constantly monitoring your data coming in, and constantly sort of adapting to, to changes, including the changes that you, you know, provide by, you know, taking some actions, super important, and I think those lessons kind of transfer well beyond finance. Derek, I’ve got a question for you.
So once you’ve identified that you’ve got the right data, how much time do you spend making sure that that data is valid and accurate? Cleansing data, if you will. Good question. We spend a ton, a ton of time. I would say maybe the, the majority of our time is actually spent on that, and I think that’s a sort of general trend in, in data science. I mean, we’re, we’re often like working with data, building machine learning models, and constantly sort of evaluating what the models are telling us.
And yeah, like I said before, if you don’t have good data, you have sort of no chance of predicting the future, no matter how powerful or fancy your model is. So
I, I would say we, in the sort of data science profession, we might spend something like eighty percent of our time taking care of the data, making sure the, you know, the data’s all there and it’s clean, and then it’s sort of free of any biases you want to avoid.
And then maybe the other like twenty seven– percent of the time would be sort of modeling and testing out different things and going from there, but a ton of time is spent on the data. I, I wanted to ask, just so we kinda set a baseline for our audience, if you could explain the difference in between, you know, machine learning, like what that means versus what, you know, the, the terminology AI, artificial intelligence, or, you know, another one that’s thrown around, large language models.
Can you kinda set the, set the stage for our listeners so that way, you know, they have some grasp of these concepts as we move forward through the conversation?
Sure. Yeah. Yeah, that’s a good, um, that’s a good question. So AI is sort of this general, uh, you know, kind of vague notion of, you know, having computers help us understand the real world and maybe predict what’s going to happen.
And- You know, how I often think about machine learning, it is a little bit different than, you know, what you might, what, you know, the people at OpenAI, the people behind ChatGPT might be thinking of where, where, you know, they’re often working with what we call natural language.
You know, they’re sort of inputting some piece of text into their model and trying to get some text out of their model, which is, like, definitely, you know, a form of machine learning.
What I sort of do more often is machine learning on the so-called tabular data sets. You know, something that you might imagine would fit in, like, an Excel spreadsheet, something like that. And, you know, essentially, we’re trying to build some, some model, some understanding of the world or even, you know, understanding of the data that we have.
And often our goal is to try to predict what’s going to happen in the future. So I, I think, you know, the general sort of learnings from data science and machine learning that people do in, in any industry are, like, totally applicable to the, uh, to the transit space.
I, I appreciate that explanation. It… You know, I, I think it’s important that we, we all kind of have that familiarity before we get too far into the, the conversation.
Uh, one thing that kind of stood out to me there, and we, we mention this a lot at, at ETA, is having good data.
Uh, is, is there a point where you’re thinking about the data set that you have where you’ve, you have that comfortability with the quality of the data set?
Um, y- you know, I don’t, I don’t know if there’s a specific point where you finally think that, okay, this is in good enough shape for me be- to be able to use, or perhaps it’s just an evolution and y- you know, you always think, well, this could be better, right? We could be cleansed more.
Maybe we’ve omitted data or a data set that we could have used and it wasn’t available to us at that time. W- what are your thoughts there, Derek?
Yeah. I would say it’s more of the latter. It’s sort of a constant evolution. Your, your data, if you have a lot of it, it’s never going to be perfect, and it’s never going to be, like, pristine and clean, and it’s never going to be like, like, e- everything you could have could ever want. And there’s going to be, you know, even in clean data, there’s going to be sort of like false signal that you, you know, you or your machine learning model might sort of catch onto and think those are what’s actually sort of causing things to, to happen and sort of changing your predictions.
But maybe, you know, these are just sort of like correlated rather than actually causal or, like, predictive. But, you know, sometimes we’ll sort of work on some data set and maybe clean it up or, you know, transform the data in some way to, to create features for your, for your machine learning model to, to learn from. And often what that looks like in practice is just sort of like doing your best, being really thorough until you start reaching some diminishing returns. And then maybe, you know, at some point you start to think, “Okay, maybe my time is better spent elsewhere,” maybe tweaking the models or, you know, training differently or something like that.
But, you know, the data is probably the most important, uh, part of machine learning as a whole. There’s one funny story, and it’s an LLM story, where there’s a, there’s a subreddit about microwaves, and it’s, it’s just like a nonsense subreddit where people just s- sometimes just make like, type in like microwave sounds, which might be like m, m, m, m, m to, you know, create sort of like the humming of like a microwave. And that data get, you know, it can sometimes get into the, like, training these large language models.
And when they go and try to train their models on that data and they see like, they see like, you know, a hundred Ms in a row, these models think like, “Okay, there’s no way we’re gonna see another M coming up again.”
And then they do, and can just sort of like poison your model. And then, you know, if you open up ChatGPT, and you type in like, “M, M,” and
ChatGPT gives you back a thousand Ms in return because that’s what it learned from that subreddit, it’s probably not what you want. So it’s, it’s a funny story of, you know, getting bad data into your model can, uh, sort of wreak havoc.
That, that, that’s very interesting. Uh, you know, uh, I want to ask more about the, the usefulness or, or why, you know, organizations need to be looking more at predictability. So, you know, public transportation today, most agencies, like with the current technology, they get to see what happened, right? Like all the historical data, they run reports, and then they take a look at it, and then they interpret the data, and from there they try to come up with strategies, right, like to make, you know, performance better to drive these numbers. So it’s, it’s a lot like reaction, right? You know, wh- when it comes to predictability, like, could you tell us like why is it valuable, you know, uh, or, or how when, uh, organizations start looking at predictability versus just like the results, like what, what’s the benefit of that? Like, why should organizations and particularly maybe transit agencies think about of like how can we start predicting versus just reacting? Yeah. That’s a good question. And one that I haven’t thought about until, until we, we talked about doing this podcast.
You know, I’ve, uh, you know, I was a sort of transit user for a lot of my life, but yeah, surprisingly, I hadn’t thought all that deeply about data science and sort of like the transit and how they, how they might interact. So, you know, this question of going from Explaining what happens to predicting the future is totally relevant, and, you know, it’s– I often come across these like the same sort of like binary of problems in, in finance.
I think the, the really important thing to think about is like what decisions could I make and what interventions could
I, could I perform, and then sort of work backwards. So, you know, maybe you’re like, you know, maybe you have some prediction where you are expecting a very high volume of like ridership in some like bus system or something like that. It’s easy to sort of like go back and explain the past or at least, you know, train some model and think that you’re, you’re explaining the past in some, some accurate way. But the real test for prediction comes to like seeing if that, if that model, if that understanding of the world keeps happening in, in, into the future.
So what this looks like in practice is like being very careful about how you train these models, where let’s say you have ten years of data. You have like historical data for, you know, maybe let’s say some like bus system and you’re looking at sort of like ridership or something like that over time.
You might take, you know, the first year of that data and train some model and then test that model the next year of data and see, okay, if I, if I had trained this model in year one, how would it have done in year two? And then after that you can, you know, take the first two years of data and see how your model performs on year three, and then sort of just like expand from that. And in the data science industry, we call that technique walk forward cross-validation, a very fancy term for like making sure you’re not cheating and, you know, making sure your model isn’t looking into the future when you, when you build it. And doing this gives you lots of sort of measures.
How good ha- would my model have been in the past? And that gives you some idea of can I use this in the future for predicting what’s going to happen. So if I understood you correctly, Derek, you, you’re talking about using that historical data because it, it’s historical, it’s already past. We know how the agency performed or how ridership was. So it’s, it’s not like, as you put it, it’s not cheating.
It’s not thinking this could happen in future years one, two, three, four. It’s this is what did happen ten years ago or, you know, five years ago, something like that.
Exactly. So Derek, once you’ve got that model dialed in, right, you’re, you’re testing against known data sets and everything’s looking good, and you’re feeling pretty good. What blows up? Good question.
So, you know, the real, real test is like you put this into production, and you start acting on it, and you, and you see how it goes. And so for, you know, for us in finance, we, you know, we can do all sorts of historical training and testing and stuff like that. But really, you know, where the proof comes in that you might actually have something is, you know, putting that into production and like actually acting on it. So for us, that means like actively trading and buying and selling stuff and seeing if we’re going to make money.
So yeah, that’s, that’s really the important thing. And where that might fail,
COVID is a good example where I’m, I’m guessing, you know, same thing happened, to you guys. If you had some predictive model for, you know, twenty twenty what sort of ridership we might expect, you might be disappointed in twenty twenty when the world kind of stopped.
So yeah, events that haven’t been captured in your historical data, you know, nearly impossible to, to deal with. And what’s really important is sort of monitoring those things live so that you can adapt as needed.
And sometimes it’s easier said than done. So let’s say in a circumstance where you don’t have an aberration like COVID, we’ve got reams of historical data. We’re adding to that data set as time moves forward. In even the most normal of circumstances, how often do you see the results over time differ from, you know, even the live testing you had done just a few months ago? Do, do you see systemic changes or, or are they more often than not, in the normal case, small changes in the results of the prediction?
Yeah. We– Depending on what we’re predicting, you know, some things are much, more stable than others. But totally, we can see, you know, something that, you know, worked in the past just stops working for whatever reason or, you know, maybe there are some like regime shift, you know. In markets maybe that, like you could think of like interest rates or, or, you know, something like that sort of changing the sort of fundamentals of how everything operates.
And, you know, the way to counteract those sorts of things is to, to get that data into your models when you can. That way, you know, if you have historical data where that regime did shift and, you know, maybe it shifts back or something like that, then you can learn from it and be in a much better place when sort of trying to predict the future. But in practice, we see this thing, or this sort of thing all the time.
And so, you know, it’s one kind of frustrating thing, but also kind of an exciting thing about, about, you know, working in finance or doing data science in general, is that often you have to just like stay on your toes and keep adapting as things change, and you move forward.
So, so there’s a portion that can be predicted based on the information and normal operating conditions, right? So I’m thinking and, and
I’m kinda like connecting this with, with transit and, and the things that we measured or, or the things that are important for transit. So a lot of times, one of the most important metrics for service quality is on-time performance, right? Like the, the reliability of the service. You know, you were talking about ridership, which is kinda like, you know, one metric to take a look at when it comes to service consume. So, you know, public transit agencies wanna be on time and, and, you know, wanna have a lot of riders, a lot of people using it. Historical data under regular conditions could provide prediction, but it, it’s kinda based on those patterns, right? Like, the, the patterns that you see in the historical data. Meaning, like typical in transit agencies, the month of March and October, they tend to be the highest months for ridership, and then you see a decline of ridership during the summer because, you know, a lot of the ridership is driven by school, you know, universities, you know, high school, so on. So there’s these kids that ride the system. So, so if you look at the historical data and feed it into a model, and you see that pattern every year, I would imagine that it’s kinda easy to predict what’s gonna happen, you know, with, let’s say, ridership the very next year, just because of, you know, that, that, that behavior. But then, of course, there’s other factors that maybe change the normal condition. So let’s say you’re adding service or you’re decreasing service or, you know, like there’s, there’s all these different external factors. So how do you go or, or what’s the, the, the difference or, you know, between just prediction based on patterns or, you know, just prediction based on, I don’t know, just, just other data that maybe needs to be injected? Like, what, what’s kinda like the difference between those two?
Sure. So yeah, a f- a few good things there. And I think our worlds are actually very similar in a lot of ways, where like when you, when you build a machine learning model to say, you know, predict something in the future, that model learns the patterns that are in the data that it was trained on. Mm-hmm. And you’re sort of like implicitly hoping that those patterns remain so that your model can still sort of like know how things go and predict things accurately.
But, you know, you brought up a good point of like, you know, let’s say your model says something, so you intervene. As a, you know, your model says there’s going to be some like equipment failure, maybe, you know, buses or trains or something like that will sort of start to fail, and they’ll need maintenance or something like that. And so then you sort of schedule more routine maintenance, and then you’re sort of, you’re sort of intervening in a way that is going to change the, the world in the future, where, you know, maybe you’re fixing up these buses more often, and so they fail less often.
And then your machine learning model might not know about that because that maybe is not, it’s not in the dataset that you’re intervening. Then your model will sort of start to fail. Your model might say, “Oh, I was expecting these to, you know, need maintenance or like start, like, you know, fail, you know, once every six months or so, but now I’m seeing, you know, the model’s saying every six months, but you know, these things are like never failing.”
And it could totally be seen as like a failure of your predictions, your machine learning model or whatever. But, you know, hopefully everybody sees that as like a net positive. You’re like, “Okay, I intervened. I made things better. And, you know, my model is not working anymore, but that’s okay because things are much better. I intervened, and, and things are good.”
Yeah, we have s- very similar dynamics in, in, in finance where, you know, us intervening can sort of mess up the model in, in some ways. So we try our best to think about how our intervention has like changed, changed the data in the past and how is it going to, you know, change things going forward, and that’s really tricky.
It’s a tough problem. I was thinking last night about, about how like important what you guys do is, where like, you know, let’s say you have some system that or some city that, you know, lots of people depend on public transportation. There’s all sorts of like cascading effects that, you know, things are late and people don’t know about it or, you know, things break down and, you know, maybe like somebody’s waiting for a bus and didn’t know it was gonna be late, and they got late to a job interview. Like, it just totally changed the trajectory of their, like, life.
Or, you know, a bunch of hospital workers need to get to the hospital on time, but the train stopped or something like that and, and then service at the hospital starts to decay because, you know, people have been working long shifts and it, it’s, um… Yeah, there’s sort of a whole cascade of things where like a lot of people depend on these services, and the more information you can get to them, probably the better. Yeah. It’s interesting how important what you guys do is. Today’s episode is brought to you by
ETA. Transit agencies rely on dozens of systems to keep service moving, but the information inside them often remains disconnected. Transit OS connects operational technology, enterprise software, and public data so agencies can bring more of their information into a shared operational context.
Transit GPT gives transit professionals a practical way to use that connected data, asking questions in plain language and getting answers grounded in agency information.
Together, they help agencies see service clearly, respond faster, and deliver more reliable trips for riders. Learn more at etatransit.com.
Yeah, I’m kind of wondering generally what sort of information transit organizations
Has currently, and what sort, what the infrastructure looks like. I mean, I’m guessing, like, a lot of people have dashboards, and they monitor things and, you know-
Yeah … there’s data that sort of lives behind those. But I’m guessing they usually don’t interact with, like, raw data. And y- what is it like, what, what, what sort of tools and data do, do people actually have to actually work with? Yeah. Th- this is a good one. It’s a, it’s a tricky wicket.
So, so Levi used the term CAD/AVL vendor, right? That’s what we’re called in the industry. It’s computer-aided dispatch, automatic vehicle location. And it all starts with really a distributed computing system on the bus, right? So we put a computer on the vehicle, right? At a minimum, it’s got cellular or wireless communications and location, often GPS.
But then you might plug that into the engine control module or the, you know, digital signage system or the fare collection system or the, you know, uh, devices that we install overhead the doorway to count people as they get on and off.
Well, unlike, you know, really deterministic data, which is, I suspect what you’re getting, right, on stock trades, you know, somebody could have nicked the wire that’s talking to the automatic passenger counter.
We’ve had maintenance systems, you know, follow our work and accidentally remove fuses. So data quality, you know, from, from a distributed system on a moving vehicle is really challenging.
Uh, and then, of course, you’ve got, you know, systems for HR and maintenance and workforce management, and then the tools they use to actually plan these networks.
And, and as Christian said earlier, you know, these are often provided by different vendors with different commercial interest, and were never built to work to-to-to-together.
And, and that’s the first problem we’re trying to solve is let’s unify and normalize this data so that, you know, AI can make good use of it.
Yeah. It’s, um, you know, thinking about it as you were talking, like, initially, it sounds like a lot of, like, mess.
Uh, you know, lots of different data streams and messy data coming in from each one and, you know, like, some go to this platform or this software and some go here and, like…
But I think also with all of that mess probably comes a lot of, like, a lot of opportunity, you know. Oh, yeah. A lot of interesting data that could be utilized and really help out in, you know, predicting the future in general.
But a problem, big problem to solve. So when I think about the ROI on an investment, some problems can probably still be solved with simple heuristics, right? And so for our listeners, it-it’s really simple, hard-coded if-then, right?
Versus machine learning, which is a lot more black box. It’s kinda hard to explain why, why things have happened. But machine learning takes a lot of elbow grease and a lot of smart people, right? And, and it’s not cheap. Boy, how, how would an organization, any organization, you know, begin to reason about whether the investment of getting good at machine learning is, is worth the squeeze? Could you… Is good enough, good enough with simpler models?
Yeah. Good question, and I don’t know if there’s, like, a great answer to it. Like, a lot of those, those things are really hard to estimate.
But, you know, it’s really good… I was thinking yesterday, it’s a really good time for us to be having this type of conversation about data science and machine learning. Mostly because it has gotten a lot easier recently. You know, there are lots of these, like, really powerful, very intelligent models that, you know, companies like OpenAI and Anthropic are building. And the kind of, the barrier to entry to doing this sort of, like, machine learning stuff has gone way down, where you can just get, you, you can now get, like, a generally smart person, you know, maybe they, they don’t know how to code super well and, you know, maybe they’re new to the, to the field. But, you know, what we’ve been seeing here is that using some of these new tools makes the barrier of entry and, you know, doing really sophisticated and high quality machine learning and forecasting and, you know, pr- trying to predict the future a lot easier. But there are still a lot of barriers. Like, if you don’t have the data, you don’t have the infrastructure, you don’t have, like, the systems to intervene or anything like that, that’s all still probably really hard, as you guys, I’m sure know. But the actual, like, doing data science and, you know, building machine learning models is no longer a thing where you, you know, you need some, some data science expert with, like, a decade of, of experience.
It’s, yeah, uh, like, like I was saying, the barrier to entry to doing this sort of thing is much lower and therefore a bit cheaper.
But yeah, the infrastructure and having the data and stuff, totally necessary, still not, still not easy, and guessing still not cheap. It sounds like we keep coming back to data, and it sounds like the problem is exasperated when these da-data come from lots of different sources. I’m curious, when we think about fintech or, you know, e-even contests like the ones your fearless leader has won in years past for insurance actuaries, are, are we getting data from a lot of different places, or are those data sets fairly normalized?
I would say it’s a bit of a mix, and it depends on, you know, who you talk to. You know, there are certain firms that specialize in, like, using alternative data. You know, maybe they, like, get some weird data stream from somewhere, and they’re like, “Oh, this is gold,” or, “I can capitalize on this.” What I’m more familiar with is sort of, like, well-structured data or, or maybe data that you can kind of easily transform into well-structured data.
But yeah, there are lots of different groups out there doing different things with, with very different types of data.
For example, some, you know, some hedge funds will get satellite image data and use that to try to predict, you know, stock prices or oil prices or something like that, where, you know, let’s say you have a satellite image looking down on some, some like tank of fluid, maybe it’s water, maybe it’s oil or something like that.
And you can like sometimes like analyze shadows and see like, oh, how filled that, how full that oil reservoir is and stuff like that. So it’s really amazing how creative people can get with sort of alternative forms of data.
And I remember, you know, during COVID, I was working as a, as a scientist, you know, doing physics experiments and our lab got, got shut down for, you know, COVID reasons. They didn’t want sort of people interacting. And
I pivoted temporarily to doing some COVID data science. And one really interesting data set that people started using is, it’s kind of morbid, where they were using satellite imagery over
China, sort of in like the Wuhan region. And they could do some like image recognition type of tasks. And what they were seeing was sort of dark, cloudy regions that were darker and cloudier than usual. And it was actually t- it turned out to be sort of predictive of the sort of, you know, number of COVID cases going forward, because those were from, uh, from like cremation sites.
So that’s a sort of extreme example of having some crazy alternative data that is actually really useful for like, in this case, thinking about what was actually going on in Wuhan around COVID where like, you know, maybe the Chinese government wasn’t being so transparent, but tough to hide satellite image data.
Derek, y- you know, one of the things you were mentioning is that the… When it comes to trusting your model, as long as your assumptions, initial assumptions remain true, then the predictions, once you kind of assess your, your first predictions quality, I guess, the quality of those predictions, then they should remain true. Meaning, to your point, like, you know, we were talking about ridership predictions.
If CO- COVID didn’t happen, right, like it was just a normal year, everything, you know, should have kept the same, then most likely those predictions could have been accurate.
Because, you know, again, if, if everything remains true, right? Like, like the… We’re not cutting service. We’re not significantly adding service. There’s no significant changes in the operating environment per se, then it’s easier to trust the prediction because there’s not other factors external from the data that you are basing your prediction on that you, you know, that are, is affecting such data or such predictions. But then that’s where the, the human judgment maybe comes into place is like when predictions are being generated, right? Like, you know, any kind of system that uses data to generate predictions and, and it says, like you were saying, like mechanical breakdowns, and it could be like, I don’t know, like your… You know, you’re gonna have more engine ge- uh, regenerations because, you know, your diesel particular filter is, is getting caught and usually about this mileage is gonna break.
But, so it’s, it’s almost prompting me to go and replace those or do something. But it’s like, “Oh, I already did that. Like, I just did that. We just went through a campaign and we replaced those filters.” Then that’s where, you know, you, you decide either you take action or not, right? Like based on all these other factors that your model doesn’t have, right? Like you, you were saying, you know, if the model doesn’t know that you already took action, then how would the model adjust the prediction?
So, so would you say that that’s pretty much like the, the, the premise behind it is as long as all your assumptions remain true in the moment of the prediction, you should be able to trust the prediction. I- is, is that, does that sound right or would you, would you modify that in any way? I think that’s pretty fair. Yeah. I mean, we have like the big implicit assumption of machine learning is that like you’re assuming the future is going to be somewhat like the past, and when that doesn’t happen, your model’s not going to work.
So yeah, the model will only know about the, the data that you feed it, and so your model is sort of like inherently limited, where, you know, may- I’m guessing before COVID you did not have,
“Is pandemic happening right now?” feature in your, like, data set. But you know, that’s probably the most important feature in, uh, you know, 2020, 2021. Sometimes you can, you know, realize what you’re missing out on and start building that into your model. For example, if you, you know, you start, uh, yeah, changing your, your like particulate filters in some, some engine or something like that, you could create a feature that’s like the number of days since you changed those and, you know, some, some bus or something like that. And if you have all the historical data where you’re like, “Okay, j- I changed it on this day and this day and this day,” then you, you have the option to sort of like build those features into your model in a way that’s kind of honest and w- you know, without, you know, having this like look-ahead bias where you’re kind of cheating and sort of knowing about the future before it, before it happens.
But you know, when doing data science and trying to forecast what’s going to happen in the future, yeah, you kind of have to constantly adapt and think about what you’re missing out on, like what information it’s missing out on and How you can get that data into your data set. That’s like a really valuable sort of like continuous feedback loop where you’re, you’re constantly trying to approve, improve and, and get the data that is very predictive of the future into your models.
Let me ask you this, Derek. Are there particular use cases that would be easier to predict than others? Are, are there characteristics of problems that would make, you know, predicting predictability, uh, more, more challenging than others? Like, one of the things I think about is the human element, right?
I, I imagine picking stocks thirty years ago was a little different before E-Trade. I remember as a youngster studying finance, I said, “Yeah, look at what the odd lot traders are doing, and don’t do that.” And now everybody’s an odd lot trader, it seems.
Yeah, what are some of the characteristics that would define a problem set that might be easier to predict, uh, versus a problem set that might be very difficult to predict? Good question. Okay, so I have two like clear examples in, in finance, but I am curious about what your similar examples would be in transit. Okay. Here, you know, there’s two ends of the spectrum. One might be, maybe this is sort of an intermediate, a difficulty task of like often in finance, we want to know how much trading volume there will be in the next, you know, upcoming some number of days, maybe just like tomorrow.
So, you know, if I want to like buy Google stock tomorrow, it’s really useful for me to know how much trading volume there will be in Google tomorrow. And that particular, what we call a target, that sort of thing that I’m trying to predict is often kind of nice and stable and like somewhat predictable. Like I can build a pretty good machine learning model and sometimes even like a pretty simple one to predict that variable. Just because it’s like a nice steady state thing, you know, I, I’ll know like on earnings days I’ll expect a little bit more and on sort of like boring days with no news, maybe I’ll expect a little bit less, but fairly stable. Our really hard problem in finance is predicting what we call alpha. These are like sources of like excess returns, and that’s sort of like where you make your money. If you can predict where these prices are going to go better than other people, and if you can act on it, then you could make a lot of money. But that problem is like super noisy and returns go, you know, up and down each day. There’s news coming in and out that might not be in your models, and that’s our tricky problem.
I’m curious what tricky problems and easy problems there might be in transit. Yeah.
Well, I, I think as far as the easier problem, you know, there are some consistent outputs that transit agencies can predict generally.
Ridership being one, operating hours and miles, those are relatively consistent. I mean, you’re going to have some differences between like weekdays and weekends.
You know, Saturday looks different from Friday, and then of course, Sunday looks different from Saturday even usually because there’s less service, but that’s part of what makes it predictable.
Y- those, those easier ones I think are the, the ones that stand out to me. Yeah, like even the number of vehicles that you’re operating on a particular day. Well, I mean, for me it would be on time performance.
A- a- and, and you know where I’m, where my mind is going with this conversation and, and a- as I’m listening to what you’re saying, it’s like the amount of variables that will have a significant impact on the outcome of a given, you know, measure, right? Because, you know, like ridership, there’s like kinda like behavior and then, you know, we see people, typically people that that ride public transit, a big chunk of the riders are recurring riders, like daily riders, right? A- and then they have a pattern. They go to school, they go to work, and every day they take the same train, take the same bus, they come back. So, so it’s easier to predict because again, you look at the historical and like their behavior is just, it’s, it’s, it’s part of their routine, right? But like when you go into on time performance, especially for example, for a, a fixed route system in a city, right? There’s a lot of elements that are outside of the agency’s control, uh, and even the data, right? Like not having the data meaning, you know, there’s accidents, you’re sharing the road with the, with the public, right? Like the, the, you’re using, you know, lanes of traffic that are shared.
So if there’s an accident, if, if there’s construction or things that the transit agency didn’t know of, then like the on time performance, you know, you could look at like, okay, on Tuesdays for the last few
Tuesdays, on time performance. has been an average, I don’t know, like seventy percent for the entire system, but out of the s- so you’re like, well, it will be, you know, seventy percent next Tuesday. And then it happens that maybe it’s even less because there were all these things that took place in this kinda operating environment that you didn’t know of, neither your model knew of. So that, that to me, the, the harder ones are the ones that includes factors that have direct impact on the performance or outcome of a given, you know, metric or situation that are, are outside of your control. So that, that, that’s kinda how I’m seeing it is like as long as the assumptions and like all the different variables, you have most of the data in there, for the most part, control or fully understood, then it’s easier to have a predictability that is, has higher level of confidence.
And then once you have all these different factors that affect your assumptions that are outside of your control, and then they just take place, it’s, it’s, it’s harder to have a more accurate, you know, predictability. That, that’s kinda how, how I’m seeing it. A-a-and, and on that, I did wanted to post a question which has to do with data enrichment.
And, and you were mentioning, like, okay, let’s say that you are, uh, replacing these filters on these buses and, you know, that’s why your model is saying, “Hey, this is gonna break,” but you already replaced them, so how about you inject that data into your model? Now your model knows that, so it would not be predicting that it’s gonna break when now it’s seeing that it’s, it was just repair, right? So, so you’re improving your model.
So when it comes to data e-enrichment and, and just, you know, throwing that term out there, I don’t know if I’m using it correctly, but, like, are those things or, or things that, that you look at when, when you’re evaluating a model where, where are tho-those additional sources of data that would make my model more bulletproof?
Yeah. Yeah, that’s, that’s, uh, that’s kind of huge in our space. You know, if we can find some new, some new feature or some new data set or something like that that can make our model better, that’s gold, and usually it’s gold for some limited amount of time, but still temporary gold is, is better than no gold. But, you know, often it’s like we’ll identify some data, you know. For, for you guys it might be something like days since last oil filter change or something like that, that, like, that’s going to be useful for a long time, and your model totally should know about that. And, you know, there’s this, uh, like additional sort of question of, like, okay, could that have been known sort of beforehand?
And I think sometimes the answer is yes, and sometimes the answer is no. And if the answer is no, then often you’re kind of out of luck. Like, unknown events coming up, there-there’s just no way for you or your model to sort of get ready for those.
I actually have another question for you guys, and… Well, in, in finance, we often have these kind of like asymmetries where like, okay, we’re, we’re predicting things so that we can, we can act on them.
But the sort of like risk or the penalties for if we like over-predict or under-predict are often like asymmetric.
For the example of like Google trading volume, if we, if we wanna buy like a million shares tomorrow, and we predict there’s going to be like ten million shares trading tomorrow, then it’s, it’s riskier for us to sort of like over-predict the trading volume.
When we trade, we move prices around, where like you buy something, the price goes up a little bit, and so we’re sort of impacting things.
And that depends on how much of the trading volume you are sort of participating in. And if you’re like, if you’re the majority of the trading volume, you’re gonna move the price a lot. So for us, it’s really sort of like detrimental if we think there’s going to be a lot of volume and then we trade a lot, versus if we think there’s going to be a little bit of volume and we trade sort of more conservatively, then, you know, it’s, uh, less risky.
And I’m guessing you guys have those same sort of asymmetries where, you know, in predicting like ridership, it’s, it’s… If you over-predict or under-predict, there will be sort of different risks associated with doing those two things.
Is that something that you guys consider and sort of like think about often or, or no? I can think of one just to stay on ridership, right? So if we over-predict ridership, there might be a cost impact in that we action that by dispatching another bus, let’s say.
If we under-predict, there’s less of a cost impact, but more of a service delivery impact because we don’t have enough buses out there. I don’t know if there’s asymmetry there, but the prediction will drive the over or under prediction, what will drive the, the impact of doing so. Can you guys think of any asymmetries?
Yeah, I’m struggling with the asymmetries part of it, but I… A-a-and this is maybe tangential. It, it’s a bit quantitative and qualitative, but in my experience of trying to plan these long range projects and say that, uh, uh, you know, this particular bus route o-on this BRT line is going to have X number of passengers in ten years, and if you’re overestimating the, the unlinked passenger trips, it’s called, then y-you, you can taint the project in a way that makes people think that the project was a failure, uh, rather than a success because it didn’t reach the target of what you predicted.
And there are some tools out there on the market that y- like Tbest is, is one that’s supposed to be a ridership predictor, uh, but it seems to overestimate a lot, and then people put that in their long range plans.
A-and, and that ends up kind of swaying some of the county commissioners perhaps, or some of your leadership, like this is going to be very successful, and it just doesn’t always turn out that way, so. It’s, it’s kind of interesting. So you guys have on your end sort of various sort of costs to consider, like one cost in dollars. You know, you over-predict volumes and you, you know, dispatch too many buses and, you know, you actually pay in dollars for that. But then, you know, if you under-predict, you know, maybe you’re like leaving some like old lady stranded in some Chicago winter storm for like way longer than it should’ve been or something like that. And those costs have very different sort of like units. You know, one is like personal suffering, and then one is like dollars.
And so it’s probably really, really difficult to like figure out what exactly to optimize for. I also am curious about, you know, what we call sort of like forecast horizons in data science, like- How far into the future are you guys typically looking?
And does it sort of span the, like a very large range from like minutes to like decades when you’re like, you know, doing big projects or something? Yeah. On the big project front, it’s decades. You know, you would do, like, Christian is familiar with this as well. Yeah, having worked at Palm Tran, we had a, a transit development plan or a TDP. Uh, that’s looking at the next twen- sorry, the next ten years. And then if you’re, maybe you’re working at like a, an MPO, a metropolitan planning organization, or, you know, some sort of long range organization or agency, then you’re looking at twenty, thirty or more years into the future. Uh, so there are some really long horizons. Yeah. So here’s an example of something that would have a time horizon of today, let’s say.
So one of the problems in transit is that buses will bunch, right? So, so they should probably travel in some distance, right, from one another so that we don’t have excess capacity at one stop and under capacity at another. And, and this goes back to your question, what would be one of the more difficult problems to solve? This is one I really wanna sink my teeth into because there’s human element, right? Is the bus driver, bus operator, as we call them, sort of tailing the bus ahead of them so they don’t have to pick up passengers, and they can just jam out to the radio, right?
But so, so that would have a time horizon of today. Yeah. And it’s interesting. So yeah, you guys have kind of a complex system to sort of like solve for and optimize for.
Like, and the thing that you might want to optimize for, you know, I get might be kind of different from, you know, what the, the operator might want to optimize for. You know, maybe they wanna go fast and then take a longer break and, and relax a little bit more. So yeah, it’s, it’s a pretty complex system that, that you guys have to solve, but it’s probably a really fun problem to work on. I think so. It, it’s funny, your problem is way more difficult in certain respects, but the outcome you’re looking for is always the same, find me the alpha, right? I was thinking earlier about the timing of your trades, right? How, how quick can we get our order in, right? Will al- also impact what you were talking about earlier with respect to, you know, how much are we gonna move the market.
But that, that’s a really astute point. It kinda depends on what we’re trying to solve for, what success looks like. Yeah. Luckily, you know, when doing machine learning in general, there’s like, there’s a few points where you do need to some, like, human input.
And the most important one probably is the– it’s often, you know, written out as like a mathematical function, what exactly you’re trying to optimize for. And sometimes it’s easy, you know, on my side, it might be like PNL or something like that, how much money we, we expect to make, you know, some- something kind of simple like that.
But it’s often like really, really tricky to figure out. So, uh, for us, we call it like a cost or a loss function or, you know, whatever we’re trying to optimize for, we have to do it mathematically because that’s the only sort of like language these ma- certain- these machine learning models understand. And so a human has to come up with that, like, optimization function, what you’re going to optimize for. And for you guys, I’m guessing that’s like incredibly tricky.
Yeah. And it, it, it does get at the heart of how complex our, our problem is. Like if you’re optimizing for the customer experience side, which we, we haven’t talked about too much. We talked about like the agency, but there’s also how I view the system as a rider, right? Am
I having to wait a long time for the vehicle? Are the stops dirty? Are they clean? Do I feel safe waiting? Do I feel safe on board?
The- there’s, there are just three… Uh, uh, I guess there are kinda three different tranches or, you know, three different areas that you can kinda take for optimizing for the agency side, you know, the rider side, maybe even the vendor side. Perhaps there are others too, you know, the board or leadership side.
So, uh, there’s, there are a lot of competing, you know, competing priorities, I think, when it comes to, to transit and, you know, people are, are complex. It’s, it’s not reasonable all the time.
Eric, let me ask you this. You know, with, with LLMs making AI very much front and center, some of our partner transit agencies may start to give more thought to using ML to make predictions, but I suspect many of them will be very new to it. What suggestions do you have for them?
How do you get started? What does good look like? Okay. So there are some suggestions I would have that are kind of general and maybe kind of obvious, but super important.
Like you need, you need data first. If you don’t have data, you, you can’t do any of this. You can’t, you can’t build any models. You can’t forecast the future in any way.
So getting your hands on the data that you need is often sort of step, step one. Well, maybe even, maybe step zero is actually thinking about like what decisions you have to make, what interventions you might be thinking about, and then thinking about, you know, what sort of things you might want to predict. And sort of closing in on this problem from, from both directions.
One is sort of like thinking about your decisions that you have to make and sort of working your way from like decisions to predictions to like model. And then from the other side of the spectrum, you have, you have data needs. You have a lot of data needs sometimes.
And maybe, maybe that means maybe you look at the data and you find out like, okay, you need some data that you don’t currently have, and there’s no way to get the historical data for it. So you better start measuring it now and start building up those data sets because they might become like very valuable in the future. So,
I, you know, I think that might be… Yeah, I think that’s what I would say are the sort of most important things. And then, you know, taking it from there, and sort of going step by step, and, and being careful, and, you know, being careful about, you know, how you train your models and, and how you validate them. And y- one thing that I think the transit industry will be good at is sort of monitoring, monitoring things. If s- any of your, like, inputs or outputs are, are changing,
I’m guessing the transit industry is often kind of on top of those if they have their, like, live data streams or somewhat live sort of data streams up and running. But, like, building up the data infrastructure is important if you want to do any sort of data science or machine learning to predict things.
And then you can take it from there. You can… You usually start by, like, you know, deeply understanding the data that you have, and then building often a really simple model. This might be like, “Okay,
I’m going to predict my ridership tomorrow is the same as it was today.” And then layering on complexity in terms of, like, the data you feed the model or the models themselves that you try, and seeing at each step was that layer of complexity worth it. If not, maybe I stop here’ or go in some other direction.
But yeah, I think it all starts with the data. And if you don’t have the data, you, you can’t do anything. So yeah, maybe that’s what my, uh, suggestion is to, to get the data first. Yeah, I, I think that, you know, one, one of the issues, and, and I think John might have alluded to that earlier, with the transit industry, is that the data lives in different siloed systems. And, you know, based on what you said, like, in order to really have good predictability and be able to actually, you know, take action and have models that are robust and, and are giving, you know, fairly accurate, you know, high confident predictions, is that you’re, you’re bringing together as much as the data that would empower the model to make the best predictions, right? And, you know, if you were to have, like we were saying, you know, like, ridership data or, you know, like, even, like, on-time performance, right? But then you don’t have anything that is data related to your maintenance system, you know, how the vehicles, how often are breaking down, what type of breakdowns they’re having. Then when you’re, let’s say, predicting on-time performance, well, it’s gonna be very hard because if you’re not including… If you only have, like, the historical on-time performance, but you’re not including, let’s say, the maintenance data that impacts your ability to deliver service on time, then your prediction is, is gonna have kind of like a low degree of confidence.
So all that to say that if you’re able to put more systems datas together, and, and then have on top of that, that machine learning, like, right, like enriching that system with all that data, then it will have higher confidence in, in the predictions. But then, you know, I think we’re in a situation right now where the industry is realizing that having good data is very important. Being able to start kind of like combining that data or, you know, putting it together, it makes, you know, their ability to understand the data and make decisions even better. But, you know, I, I want to ask you about, you know, now starting to take action.
So what are some examples maybe you can share of when, you know, maybe some actions are automated, maybe you’re just looking at some alerts, and, and, and how do you measure the level of action, right? The level of automation with the data. Like, is, is there something where you guys in, in what you do trust the data to the level that automatically takes actions or, you know, you have layers of actions? Like, sometimes it could be just alerts or recommendations. Like, i- is there something that comes to your mind that comes to that?
Yeah, and it’s like, it’s this difficult layer after you’re… You know, let’s say you do all this complicated, or you do something like data set, you build data sets, you, you build machine learning models, and they have predictions. And like, you have those, and maybe those look good. And then you have the kind of difficult task of, you know, figuring out what to do from there.
That is often very tricky, where, you know, there’s a lot to think about. You have to think about, like, the consequences of acting versus not acting.
Mm-hmm. And the cost of, of, of doing those things. And like, I don’t know. You, um, let’s say your, like, your machine l- learning model has some prediction, and this has like these, you know, these, these buses are gonna be really late. Like, should we send an alert out to everybody’s, like, everybody’s phones, you know, that have some app or something like that, and let them know?
Or am I maybe wrong? Will I be like, will I be spamming them? Will this happen all the time? Will I be constantly alerting people? So, you know, I think it’s probably tricky, and probably really, really case dependent, and depends strongly on the consequences of acting or not acting.
But one general thing, like, you know, as people are considering building these systems and sort of acting on them, is to, like, probably start slow. Take, take baby steps. Sort of like, you know, you build something, and let’s say you’re, you’re thinking about having it send out some automated alert. Looking back in history and seeing, like, how many times would I h- would I have been alerting people, or how many times would I have been acting, you know, changing these oil filters, whatever, really relevant.
So, you know, paying close attention to, to, you know, what would have happened in the past if you acted on this in some automated way. Often what it means is you can sort of, like, tweak a threshold.
You know, maybe there are some, you have some confidence interval and, you know, maybe you only want to alert somebody when you’re, like, 90% or more confident or something like that. You can often, like, identify these knobs that you can tune to sort of vary how often you’re acting on things to give you a reasonable, like, density of actions. The more knobs you have like that, the more complex the sort of system can be, but the more sort of tunable it is, which is a really nice feature to have.
So yeah, a lo- as a long way of saying it’s, like, probably very case dependent, and it depends a lot on the risks of acting versus not acting. So if I’m hearing you clear, MVP1 is not an agentic agent that automatically dispatches an autonomous vehicle.
Yeah, totally. Well, I, I have one question. I know we’re running out of time, and this is an important one. Do you have a model,
Derek, that’s got really high confidence that can help me predict the over under for Miami’s, uh, game against Kansas City this weekend? I wish. Yeah.
I’ll, I’ll build one and I’ll text you the, text you the predictions. Thank you so much, Derek, John, for joining us on another episode, of Stop Requested. And to our listeners, thank you for tuning in as well. Your support means so much to us. We’ll be back next Monday with another episode of Stop Requested.