On Rails
On Rails invites Rails developers to share real-world technical challenges and solutions, architectural decisions, and lessons learned while building with Rails. Through technical deep-dives and retrospectives with experienced engineers in the Rails community, we explore the strategies behind building and scaling Rails applications.
Hosted by Robby Russell of Planet Argon and produced by the Rails Foundation.
On Rails
Eddie Galindo & Kagen Hearn: Rate Limiting For Your Customers
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
Eddie Galindo is an engineering lead at Ascend and Kagen Hearn is a senior software engineer there, working on a Rails app that helps insurance agencies move and reconcile money.
Robby chats with Eddie and Kagen about why Ascend still runs one Rails monolith five years in, what changed when much larger customers started arriving, and how the team went about finding the real limits of their data ingestion pipeline. They also get into feature flags versus database configuration, and where AI tooling fits in when the software you ship moves money.
Links:
Ascend - https://www.useascend.com/
On Rails is a podcast focused on real-world technical decision-making, exploring how teams are scaling, architecting, and solving complex challenges with Rails.
On Rails is brought to you by The Rails Foundation, and hosted by Robby Russell of Planet Argon, a consultancy that helps teams modernize their Ruby on Rails applications.
[00:00:05.12] - Robby
Welcome to On Rails, the podcast where we dig into the technical decisions behind building and maintaining production Ruby on Rails apps. And I'm your host, Robby Russell, and I run Planet Argon. For over 21 years, we've helped teams maintain and evolve their long-lived Rails apps, so I tend to approach these conversations through that lens. In this episode, we're joined by Eddie Galindo, who's an engineering lead at Ascend, and Kagen Hearn, a senior software engineer at Ascend. Ascend builds software that helps insurance agencies automate the movement and reconciliation of money. Today, we're going to explore how their team has evolved the Rails app from its earliest days into a platform supporting multiple products, increasingly large customers, and billions of background jobs, all while continuing to invest in a shared Rails monolith. Eddie and Kagen join us from San Francisco, California. All right, check for your belongings. All aboard. Eddie and Kagen, welcome to On Rails.
[00:00:57.18] - Eddie
Hey Robby, happy to be here.
[00:00:59.06] - Kagen
Hey Robby, thanks for having us. Glad to be here.
[00:01:01.19] - Robby
We've had a chance to talk a little bit ahead of time, you know, over the last couple months or two and prepping for this conversation. So I'm really excited to kind of dig into things. So before we get into some of the topics I wanted to bring you on, first question I want to ask each of you is what keeps you on Rails? How about you, Eddie, first?
[00:01:17.23] - Eddie
Yeah. Um, so I think for me, it was actually the first framework I used. In the last year of college and the, I think, right after graduating. So it's been sort of part of my whole career. I stopped using it for a few years and it was nice coming back and like having it be the same. I think I kind of bought into the philosophy on like Ruby on Rails and all of that. So I think it's just been part of like how I've been building things forever. So it's like, A lot of the principles kind of resonate with, with me.
[00:01:51.12] - Robby
What about you, Kagen?
[00:01:53.09] - Kagen
Yeah, for me, it's a more recent introduction to my career. I started working with it in my last position at Middesk, but I really fell in love with it. I really appreciate how much Ruby on Rails and Ruby as a language lean into their own design principles rather than trying to hedge against them. It's very much a dynamic object-oriented programming language and it doesn't try to be anything else. And because of that, it lends a lot of expressiveness to the language and very elegant designs.
[00:02:29.10] - Robby
And out of curiosity, Kagen, what did you work with prior to getting introduced to Ruby on Rails?
[00:02:36.06] - Kagen
Yeah, I've worked with a bunch of things. I've done Node.js on the backend. I've done Python and Flask on the backend. And I had a little bit of a phase where I was doing some data engineering work. Which involved Python, Scala, a bit of SQL. Oh, interesting.
[00:02:51.16] - Robby
And Eddie, you mentioned you had started off in Rails and then you kind of went away for a little bit to do some other things. What were you, where did you wander off to in the tech programming language world?
[00:03:05.02] - Eddie
So I joined the big tech world. I used to work at LinkedIn and I was primarily just doing front-end work there. And then from there, another startup that I joined right after, it was just a Node.js shop. So it was all JavaScript and stuff.
[00:03:20.19] - Robby
What framework, front-end frameworks were you working with over at LinkedIn? Was that like Ember there or?
[00:03:26.16] - Kagen
I know that they used to—
[00:03:27.04] - Eddie
It was a lot of stuff. So when I first joined, they were actually still in YUI. And then from there it was like a little bit of jQuery and then we migrated to Ember.js. And then by the time I was leaving, I think there was also like some conversations about like going over to, uh, React. So it was a little bit of everything.
[00:03:44.03] - Kagen
Okay.
[00:03:44.13] - Robby
I think I knew a few people that were kind of wandered over to LinkedIn at some point with the, with the Ember, with the Ember crew that I think went over there at the time. So anyhow, so that's a conversation for another day. So like, uh, I got a lot of things I want to talk with both of you. So, but before we get into those things, I feel like we should probably provide a little context for our listeners for what is Ascend? Like, what is this? Maybe, uh, Eddie, you can give listeners a quick overview of Ascend and some of the problems that you're solving for your customers?
[00:04:11.21] - Eddie
Yeah, so Ascend is a company that started around 5 years ago with the mission of building a financial platform for the insurance industry. So we specialize on all like money-related problems specific to the insurance world. So we primarily work with like insurance agencies across the country trying to solve primarily 2 money-related issues. We focus a lot on like, on like small and medium-sized agencies, helping them collect money, disburse those funds. And then on more on the enterprise side, we kind of facilitate workflows for them to operate their bank accounts and just money reconciliation when like commissions that they're receiving and stuff. So a lot of it, the big focus of the application is around money and how can we make you more efficient when it comes to dealing with those type of problems.
[00:04:58.13] - Robby
What's Ascend's like business model? Is it like a broker in between that or? How does Ascend make money with that being part of that equation?
[00:05:06.15] - Eddie
So it's a few channels. So one is we're a SaaS platform, so that's a part of it. On our transaction product, it's really just some fee on our financing book. So like we help people also to finance the policies that they're purchasing. So that's part of it. And more on the money transactions side of things too. So we have a few different channels.
[00:05:31.15] - Robby
Roughly how large of an engineering team is there at this point in time?
[00:05:36.22] - Eddie
Today we have around 20, which is probably the largest it's been.
[00:05:40.23] - Robby
Okay, thanks. And you said there's about 5 years into the development now as an application. So Rails new was 5, 6 years ago or so?
[00:05:50.08] - Eddie
Yeah, it's been Rails since the start and we haven't moved away since.
[00:05:54.05] - Robby
Were any of you there around that period of time or how? Soon after, did you join?
[00:06:00.10] - Eddie
I've been there since the start, yeah.
[00:06:02.13] - Robby
What about you, Kagen?
[00:06:03.16] - Kagen
So yeah, I'm fairly new. I joined the team just a few months ago, right as our SaaS side of the product is like hitting its hockey stick in terms of customer growth.
[00:06:14.23] - Eddie
Interesting.
[00:06:16.01] - Robby
So it's kind of fun to get the two different perspectives for our audience here. You got someone that's been at the beginning and someone that's kind of new. So in a few months, how, Kagen, maybe if I recall from our conversation, you work more on the direct bill side of things. So Anything you can add to help kind of further paint that picture for, or for our audience here?
[00:06:32.13] - Kagen
Yeah, so the SaaS side of the product is mostly tailored towards the accountants that work at the insurance agencies. So for direct bill, the insurance carriers that underwrite the policies will collect the premium and then they will pass the commission for that policy along to the insurance agency afterwards. They will receive the deposit for the commission asynchronously, um, relative to like, and then they'll get a statement for it either before or after they receive the money. And they need to reconcile where all of the money comes from, from their deposits versus, um, what's in all of the statements and do some matching. And the direct bill reconciliation team helps with that.
[00:07:13.05] - Robby
Are either of you allowed to, or are you, are you allowed to share any terms of like scale, like how many, how much you're processing on a regular basis in terms of Number of transactions are a little off the books at the moment for us.
[00:07:28.11] - Eddie
I think the money stuff we probably beat, but I think on numbers of statements, probably Kagen can mention a bit on that.
[00:07:37.17] - Kagen
Yeah, we are processing more recently like tens of thousands of statements in the past few months, and some of those statements can have tens of thousands of entries on them also.
[00:07:50.14] - Robby
Oh, wow, okay.
[00:07:51.05] - Kagen
That need to be reconciled. So I think the number just of, um, statement entries that we've reconciled has more than doubled just in the past few months over the, this period where we're onboarding lots of new customers. And that number is going to continue to grow, um, pretty dramatically, many more multiples, um, as those customers ramp up because a lot of them are just getting started and there are also more in the pipeline. So. Primary scaling factor is number of statement entries that we're ingesting or that we're processing. And that is growing by the millions.
[00:08:27.16] - Robby
Okay. Well, I'm, I'm, I know we're going to want to dive into some of those, those kind of fun scaling challenges that might be navigating. So I'm curious before we kind of get into that, since Eddie, you were kind of around at the beginning, do you recall why Rails was considered? For this particular platform?
[00:08:46.16] - Eddie
Yeah. So I think, I mean, a big factor of it, one of the co-founders used to work at Instacart. He had Rails experience. I was like, oh, what's the easiest way to get started? So he bootstrapped the app and that was kind of it. And then from there, I think a few of the attributes from Rails really helped us at the beginning. It was like, he's just a speed tube, like actually having a product out. Like he was really easy to kind of get the initial thing going, not having to spend time making 100 decisions. So like, oh, what am I going to use here? What am I going to use over here? So it was just really like, hey, what's the fastest way that we can put something out in front of customers and make sure that we're going down the right track and focus more on the product side of things rather than like making 100 technical decisions that might not be relevant if we don't actually survive. Sure.
[00:09:35.13] - Robby
So it was a little bit of a familiarity thing, you know, or, and so there must've been probably, probably it was more, maybe more than just a familiarity thing, but were you recruited because you specifically, you had worked with Rails and they were trying to bring in people that had worked with Rails before, or how was that kind of, what did that look like? I'm just thinking for anyone listening that might be thinking of starting a new startup venture and they're like, do we actually need Ruby on Rails developers to be able to build this Ruby on Rails application and platform? Um, it's something I talked about with a lot of potential business owners or budding business owners, but I'm kind of like, what's your take on that? Do you think that's true? Like you need to have people with experience early on or do you think you can kind of like let people from other skill sets kind of come in and pick it up pretty quickly?
[00:10:19.03] - Eddie
I don't think necessarily. If anything, actually at the beginning, it was probably like 30% of the team that actually knew Rails. It was more of like, hey, we used to work together at the previous startup. It was more of a people thing of like, oh, we like working with each other. So we'll continue working out together. And then from there, I think Rails for the people that hadn't used it, it was like easy to pick up. I think we've also tried to not like deviate from like your typical Rails app. So we like relied heavily on the guides and just like stuff that was out there, right? I think there's a lot of content that people can look up to kind of. Familiarize themselves with the framework.
[00:11:01.04] - Robby
Sure. Kagen, since you joined more recently and coming to an established Rails application, what was your onboarding experience look like? Does it, did it look and feel like a pretty typical Rails application from your perspective?
[00:11:12.19] - Kagen
Yeah, it really did. I was pretty impressed by actually by how well Rails conventions have generally been followed. It was pretty easy to onboard whenever wherever I expected to find things, I would generally find them. Um, so I do think that we have benefited from staying mostly on Rails conventions over time, especially if you're somebody that's already familiar with Rails and you're onboarding.
[00:11:39.09] - Robby
Okay, that kind of makes a lot of sense. I'm curious, uh, Kagen, like, how much do you feel like the Rails guides has provided— you mentioned like it was really helpful, and you too, Eddie. In terms of like, do you feel like there's aspects to the Rails guides that your team was able to kind of quickly pick up on most things? Or where do you, how does your team kind of like decide like when you're going to do something like the Rails way versus like coming up with your own innovative solution? Because we know that Rails doesn't provide all the solutions necessarily that an organization like yours might need.
[00:12:12.16] - Eddie
Yeah, I think I can speak on kind of more of the things from the start. So I think we, we use the guides as a way of like, hey, somebody probably already thought through like, oh, how do you do this? So it's like, this is the place where we start. And we've actually haven't run into too many situations where it's like, oh, this doesn't exist on Rails. If anything, like, hey, you can like very easily find a gem. But I would say like maybe 80, 85% of the work, like, hey, you can just go to the guides and see how something is done. Like, In terms of the application, I think the interesting pieces are on figuring out the product side of things rather than like, oh, how do you validations? How do you send a JSON response back? Those things, we've been doing it the same way and we don't have to keep making that decision. So I would say the majority of things are kind of covered there.
[00:13:08.17] - Robby
Is this platform primarily a monolith? I know you work on kind of like separate teams, so to speak, within the same engineering organization. What's that kind of structure look like today?
[00:13:18.12] - Eddie
So we do have a monolith application for our, what we call, like, our core platform. One, because the product is meant to be a platform and you have customers kind of using various pieces of it. That has kind of kept us on that monolith route, and it has also been, like, way easier to manage than having multiple microservices.
[00:13:37.19] - Robby
I'm curious because you mentioned with Kagen coming in, it's kind of working on a different part of the platform, and you mentioned, I guess, maybe it was a SaaS part of the platform. Was there conversations around splitting that up and having them separate, or is there enough shared? What's the kind of rationale for not separating it? For anyone listening, I'm like, well, there seem like different teams. Why wouldn't you?
[00:13:58.11] - Eddie
At the foundational level, like some of our modeling that we do, we do share a lot of concepts. So for example, we integrate with external systems and we run like some normalization of data there. And actually the whole product benefits from that. Aside from that, there is like other core components like your user management, your organization management and all of that that are still shared. I think a big reason why they still make sense to be in the same place. And at the same time, I think it's also a lot of the infrastructure side of things. Like it's been way easier to manage. And I think for a small team, that's like pretty important, especially in the world that we're in and the money side of things, but how we have to be compliant and all of that, there's not 100 applications that we have to keep monitoring closely. So that part is also important. It's just like, it's easier to do. And I think we'll try to push it as far as we can. Like once it gets hard to manage, we'll probably think about it again.
[00:14:55.14] - Robby
Sure. I'm sure it'll come up again. But how does your team kind of establish boundaries at this phase of it? Because you're still kind of 20 people and maybe, I don't know if you're going to be hiring more people in the near future or not, but like, How do you kind of think about boundaries of like, which part of those two teams is responsible for what, or is it kind of a little bit implicit?
[00:15:14.15] - Eddie
Yeah, so at the product level, like each team kind of owns their product, like, and they're responsible for that. More of the platform side of things and like all things that are shared across is really more of a team effort. Or like, hey, we all kind of own it together and are responsible for keeping that clean. I think that eventually, as we add more people, it might break, but I think for now it has worked well for us.
[00:15:43.22] - Robby
What about for you, Kagen, as someone that's been primarily working within the direct bill side of things? Does the shared monolith feel like a superpower or a constraint at this point?
[00:15:54.11] - Kagen
I think it provides a lot more benefits right now than it does constraints. In my opinion, a lot of the benefits of microservices is organizational. It makes it easier to deploy changes, um, to a specific team's domain, um, independently without interfering with anything that other teams depend on, right? Like you can isolate dependencies, um, and deploy independently for each team's domain. But I think right now we're still at the point where we share a lot of the models, right? Like a lot of things in the insurance industry are common regardless of whether it's agency bill or direct bill or whatever other accounting task. That you have going on, such as the policies, any lines of business, clients, etc. And all parts of the shop share those models and continue to benefit from being able to add to them in one place rather than being able or needing to update them everywhere anytime you want to extend functionality.
[00:16:51.01] - Robby
Will the direct bill team be growing more as this— you've mentioned this hockey stick growth type of thing, and like, is that part of like the plans right now? Are you trying to stay kind of in a realm of— and how's the team divvied up? Is it like 50/50 approximately at this point, or is it just a few people working on one side of it?
[00:17:07.11] - Eddie
So we're actually like split up into like 5 different teams today for like 5 kind of products that we have in the market. And the teams are pretty like, I would say, like evenly split. There might be some that have less people right now just because they're newer products, but They're pretty like evenly split today. Yeah, I think in terms of like, we're looking forward to like continue to grow in each one because we usually like start up most of these products are just like tackling the surface. And I think as you like go on and like discover new things, like there's always areas of more depth that you can go into. Sure.
[00:17:46.12] - Robby
That makes a lot of sense. So Kagen, you know, in our prior conversation, you had mentioned that, you know, you'd already talked about how you've been onboarding new customers on the direct pay side of things. So, uh, maybe some of them were dramatically larger customers. And so what sort of impact did that have on the platform that your team had started building already? And how much larger relative to prior customers were— are we talking?
[00:18:11.16] - Kagen
Yeah. Um, so on the direct bill side of the house, up until recently we had, um, 2 relatively large organizations using the product. And the product was fairly hastily put together. I think, Eddie, what was it, about a year ago or so, 1 or 2 years?
[00:18:35.06] - Robby
Are you blaming Eddie? Yes.
[00:18:37.22] - Kagen
No, of course not.
[00:18:39.09] - Robby
Don't be so hasty, Eddie.
[00:18:43.03] - Kagen
But There's this interesting dynamic going on where we're trying to ship at startup speeds for enterprise, large enterprise customers. And as we've started to onboard lots of customers, some of the prior assumptions that were made started to break down under the pressure, especially ingestion scaling. So in order to service direct bill customers, we need to ingest their policies, their lines of business, um, any of their clients, et cetera, from the, um, agency management system or AMS that they work with. At the time that we started getting this, um, deluge of new customers, we were able to process maybe like 500,000 to a million new policies in a day. And the limit was just latency. It's a lot of external requests that we have to make. They take a lot of time and each one holds like a Sidekiq thread. We use Sidekiq for asynchronous processing and each one holds a Sidekiq thread the entire time that it's waiting on the request. And so there's just a strict limit to with how many Sidekiq workers we had, how much volume we could process in a given day. And so we needed to start driving that up.
[00:20:03.04] - Robby
Out of curiosity, are the, you say, what's an AMS for anyone out there listening that doesn't know what that is?
[00:20:07.12] - Kagen
Yeah, it's an agency management system.
[00:20:09.09] - Robby
Okay. Um, and are these more modern platforms that you're integrating with, or are these older?
[00:20:16.10] - Kagen
Most of them have been around for a while.
[00:20:18.11] - Robby
Okay. We're talking as old as like XML, SOAP services.
[00:20:23.07] - Eddie
Yes.
[00:20:23.15] - Robby
Okay.
[00:20:23.23] - Kagen
So SOAP services, XML, all that.
[00:20:26.12] - Robby
Okay. Um, any CSV transfers over FTP or anything like that, or?
[00:20:33.10] - Kagen
I, I don't think we're ingesting anything that way.
[00:20:36.01] - Robby
Okay. Then you're not going too far back in the, in the, the eras of data transfer between different platforms. But, uh, so yeah, so you got all these background jobs using Sidekiq and they're happening asynchronously. Well, can't you just add more workers? Like what's, what's the, what's the solution there? Like, I mean, when you say you couldn't scale up more, is it just that there can only handle so many in a day?
[00:20:59.08] - Eddie
Yep.
[00:21:00.01] - Kagen
So when we first started facing this problem, We weren't sure actually what the scaling boundaries were because we had never tried to ingest more as we had, had never really needed to. And so the first thing we needed to do was try to understand what the actual scaling limiters were. So during off-business hours, we would just kind of turn the knob up on the number of workers to see what would happen. And the first limitation that we noticed was just the number of database connections. Um, that those would check out. Sidekiq will take a database connection for each thread the first time that it does any database work in that thread. And all of these workers are doing database work. And so we would start pushing pretty close to Heroku Postgres's connection limit pretty quickly when we started trying to up the number of Sidekiq workers.
[00:21:51.12] - Robby
Are you still hosting this on Heroku and postgreSQL right now.? And did you come up with any interesting solutions to that, or is it something that you just had to like figure out how to throttle at this point in time?
[00:22:04.20] - Kagen
Yes. Um, there's the standard solution to connection limits is just, uh, transaction pooling. So we brought in PgBouncer for client side as a sidecar for the Sidekiq workers specifically that are doing ingestion. We didn't apply it to all of our workers because we weren't, in order to implement transaction pooling, there are some Postgres features that you actually can't use and it can be a bit of a heavy lift to actually like comb through the application and make sure that none of those places are currently relying on it. So we introduced it in a limited capacity for ingestion workers, which really helped us out.
[00:22:43.18] - Robby
When you're processing all this and like say, I'm trying to just, walk through the workflows. Let's say someone you're ingesting maybe for a customer and going through their, their things. Are you able to then run another one for another customer in parallel, and, or is it then has to wait for the other one? Like, what's that? Have you come up with some— what, how is that currently looking, or how did it originally look, and how have you kind of been evolving and, and morphing that?
[00:23:10.06] - Kagen
Yeah, the way that it looks is, um, We have a cron job that runs every 10 minutes that will poll each of our customers' APIs for any change events. So this will tell us which policies have changed in the last X time. The window is configurable, how far you want to look. And for any of those policies that we find that have change events, we'll create a record in our Postgres database to represent that change event. Then we have another scheduled Sidekiq worker that runs on a cron that will pick up change events and then push them into the respective processing queues. The policy sync workers, for example, this is the primary one, will be enqueued by this worker that picks up change events. That allows us to limit the number of change events that we process at a given time. That's one lever to limit concurrency, is how many change events this worker will pick up at a time. So we can stack up a backlog of change events and then process them as quickly as we're able to or as we need to. And then the policy sync workers will pick up those change events.
[00:24:19.20] - Kagen
Those have some of their own side effects. Those will sync transactions that are related to those policies, for example, as a secondary side effect. But all of the customers can run on the same shared resources at the same time.
[00:24:33.01] - Robby
When it comes to like that workflow, out of curiosity, if you're checking every, polling every 10 minutes, what if the policy changes 12 minutes later? Are you having to cycle through? Is there anything about those policies that you're having to then go back and check things as you're actually processing that policy again to like double-check the data? Or is it kind of, do you poll and just know that and then you fetch it and you handle that transaction then? And then it's, it's, it's safe enough that it can keep running 12 minutes later and it's not going to be an issue.
[00:25:04.06] - Kagen
It's, yeah, it's completely fine for it to run 12 minutes later. We, the window that we look back is not 10 minutes, we look back further than that to give us some buffer room to make sure that we're not missing policies or missing any change events.
[00:25:17.07] - Eddie
Sure, sure.
[00:25:17.23] - Kagen
We have a daily sync that will look back over the past day, and then over the course of the day, we also every 10 minutes will look back at policy change events from, I'm not sure what the latest is, the past hour.
[00:25:30.19] - Robby
I see. So, so you have, you kind of have a little bit of a fail backup situation happening just in case to catch things, hopefully catch things.
[00:25:38.14] - Kagen
Yes.
[00:25:39.13] - Robby
And how has that, you running all this, you know, fetching all this stuff from these different systems, has that all been super smooth? And in terms of like those platforms have been able to handle the load that you're requesting from them?
[00:25:54.00] - Kagen
Yeah, that's a great leading question. Um, Thank you. What we have observed after we solved our transaction, our database connection problem, and with transaction pooling, the next limiter that we observed is that our users' own APIs are only able to handle so much ingestion traffic that we send at them. If we send too many concurrent requests at a time for policies, we'll start getting like 504s, server failures. In the worst case, we might even cause them disruption on their own service if we're not careful.
[00:26:28.23] - Robby
How have you handled that? You just— a lot of retries, or—
[00:26:33.13] - Kagen
yeah, well, we don't want to bring the customer down, so we can't just retry it.
[00:26:36.07] - Robby
Sure.
[00:26:36.22] - Kagen
Um, so we implemented a concurrency rate limiter, and this is inspired by a Stripe blog post on rate limiting from 2017, inspired our implementation. Specifically one on concurrency rate limiting. They have a really nice GitHub gist that gives you a sample implementation, um, and we use that. And the way that works is it puts a set in Redis for each individual integration that has outgoing requests. And each time that we want to make an outbound integration request, say to fetch a policy, um, that worker has to try to obtain a slot. Obtaining a slot is just adding an item to that set. Each organization has some configuration that tells how much concurrency they're able to withstand, so how much we can send them. If the number of items in that set is equal to or greater than the limit they have configured, then the worker is unable to acquire a slot. And so it'll keep trying once every like 100 milliseconds or so, and then eventually it'll time out and it'll just re-enqueue the job.
[00:27:50.13] - Eddie
Interesting.
[00:27:50.21] - Robby
What's, what's your, what's, what's been the kind of behind the scenes? What's the, the high level? How does that, what does that look like from, uh, within your Ruby code?
[00:28:01.10] - Kagen
Yep. Um, so the one thing this forces us to do is make sure that all of the jobs that we enqueue are idempotent because a job could fail to acquire a slot at any time and it's processing. In any rate-limited task, we try to just put one request in the job, make sure it only has one side effect. These are general Sidekiq principles, but they're not always super easy to follow. And so at the beginning of a policy sync, for example, we will attempt to acquire a slot. If we're unable to acquire a slot, then we'll just re-enqueue the job at the back of the queue. This helps make sure that other organizations are able to continue processing volume even if one is currently being rate limited. And then that will enqueue some Sidekiq effect workers of its own, like syncing transactions, for example, that is also rate limited. And because all of this is implemented using a Ruby class with a context block, it's opt-in. So requests that are executed from like production user paths, for example, are not subject to the rate limiting. If we're doing something where the user is clicking buttons and that needs to make external API requests.
[00:29:08.20] - Kagen
That's not subject to rate limiting, but ingestion requests are subject to rate limiting.
[00:29:14.06] - Robby
And do you reserve a certain allocation of slots then for user-driven things then?
[00:29:21.08] - Kagen
So user-driven things actually don't— aren't subject to the slots at all because they tend not to push so much volume that they put the customer's API under stress as much as the ingestion ones do. So the slots typically are only used for ingestion.
[00:29:36.08] - Robby
Okay, interesting. And Eddie, when you're thinking about designing systems like that, how do you think about protecting? Are you using these types of functionality in other parts of the application as well?
[00:29:48.10] - Eddie
No, I think today, probably this product is the one that's pushing the boundaries on our application in terms of scalability and performance. So, kind of connected back to one of your initial questions, it is probably one of the cons of the monolith of like, hey, The different products might have different requirements. We're just living with the fact of like, hey, we need to pay attention to infra on some side of the system. Like we'll have to do it for everything. But yeah, I think this is probably the one that has been pushing it the most. And I think also another important aspect is like, it's probably things that we didn't consider at the beginning. Like I think everything that Kagen described is like, hey, It didn't start this way. We just kind of slowly got in there. Like, let me take it back to the basics. I think what he described started really as like, hey, it's just a worker that's just going to make an API call and ingest stuff. And we just kept running and running until we like kind of hit the limit.
[00:30:43.12] - Robby
So have you been able to work out with your different, with your customers and figure out what these, do they make it very obvious what the limits are of how often you can be hitting their their app platforms, or these, just trial and error, and you're like, well, we seem to be able to do this many, and then it starts to— we start getting a bunch of errors, so we'll, we'll just kind of, we'll pencil in this number right now until we see there's an issue again.
[00:31:07.13] - Kagen
Yep, most customers' limits start at a fairly safe and conservative value, um, and then if we feel like we need to push more ingestion volume than what that conservative value allows, we will assess based on the size of the customer, whether they have like on-prem infrastructure or they're using something that's shared, for example, to try to estimate what we can actually push through. And there's a few customers that are only able to take like a very limited amount of volume as well. And so we actually have to scale them down.
[00:31:43.12] - Robby
Interesting. And it's not been as easy for you to convince them, can you just increase your, your side of the equation and add some more hardware over there.
[00:31:50.21] - Kagen
Yeah. Just spend more money.
[00:31:52.04] - Eddie
Yeah.
[00:31:53.03] - Robby
Uh, you're, you're, you're trying to provide them a service, right? So you're having to do all this stuff to work around their limitations. And, uh, anyway, like I said, maybe save that for another day. I know that, uh, I think, I don't know, I can't remember which one of you were mentioning this. It might've been you, Kagen, but I think you mentioned that in the platform, uh, there's a lot of feature flag and maybe configuration complexity in the platform already. So is that, is that, is that true?
[00:32:22.21] - Kagen
That's definitely true. That has become a lot more true over the past few months as we've been onboarding all of these large customers. This is, I think, natural to like enterprise B2B contexts. Each customer wants their own version of the product. There's different things that they need. They all have their own accounting workflows. And ways that they think about reconciliation, different processes. Some of them have acquired other agencies and that affects their processes as well. And so the set of features that we expose to each customer and the configuration that we allow them is different depending on which one it is.
[00:33:04.18] - Robby
So there's a lot of conditionals in your code to do more than maybe just swap out a logo?
[00:33:10.21] - Kagen
We, we try to avoid continually adding them, but there are necessarily a lot of conditionals.
[00:33:16.17] - Eddie
Yes.
[00:33:18.07] - Robby
How do you see that scaling? Like, you're kind of— if you're months into this process right now, and I've talked to a lot of companies that will be years into the process, and they find themselves with this tangled mess of being like, it's really challenging to debug things in different environments, like, and Maybe part of that was because back in the, back when they started, they maybe didn't get to leverage a lot of things like feature flags or things like that. So are you using much feature flag type approaches to this? And is this just some big config files or how does your team kind of think about this? Thinking about the long term, but also trying to, you're obviously trying to solve problems today for the customers you're bringing in and you're still learning, I would imagine. You probably don't have a clear vision on how this would look like in 2 years at this point.
[00:34:01.17] - Kagen
That's right. We have two primary levers, two mechanisms that we can use. Um, we use feature flags pretty extensively and the idea for those is gating actual product functionality. So what views are available to who, um, who sees which buttons in the UI, stuff like that. And then we also have organization configurations, which live in our database, and this is to tweak functionality for products generally that everybody is on. Um, so this is things like for direct bill reconciliation, for example, we need to match the entries on a statement to the policy that the entry is for. And every organization has different rules for exactly how they want to match those entries. And org configurations allow us to provide a way for customers to say, we want to match it this way in this context and this way in that context and allow everybody to have it the way they need it.
[00:35:03.21] - Robby
When you say like you have things in your database, is this like one potentially growing god object at some point or?
[00:35:12.04] - Kagen
Yes, the organization configuration. Exactly. It is currently a large growing god object.
[00:35:17.15] - Robby
Are these a bunch of like attributes? Are you like actual columns or is this like a JSON? Like what's that look like behind the scenes right now? If you don't mind sharing behind the little bit of the spider web.
[00:35:30.21] - Kagen
Yeah, it's columns on a database model. It's not a JSON, thankfully. It's so anytime you want, one of the things that adds a bit of friction and makes it harder to continuously introduce new organization configurations discourages it a little bit is that you need a database migration for each organization configuration that you wanna add. And so each one is a database column on a model that belongs to the organization. I see.
[00:35:56.15] - Robby
Is this a multi-tenant application and do you keep the database, the data very separate, or is, if you're allowed to kind of share that?
[00:36:05.06] - Eddie
So organizations, like their data is split up, but we still have a single, a single database. Okay. One more thing to add on the configuration. I think we, we are thinking of actually start splitting it up more on the domain that it's configuring. Like, oh, we have payment configuration or we have our reconciliation configuration for our other products. So to kind of ease the modeling there and not having one huge object for configuration. And I think the decision for us on like, hey, actually putting it in the database because those ones are meant to actually be part of the product and how they work and they're going to stay there forever, if that's the case.
[00:36:49.01] - Robby
That's interesting. How do you decide when to use a feature flag, maybe putting something in configuration? And what are you using for feature flags at Acuras?
[00:36:58.05] - Eddie
For feature flags, we're using LaunchDarkly today. Okay. So yeah, I think it's kind of that principle, like, hey, whatever is meant to be short-lived and like, hey, as we're working with more enterprise customers, some things might require approval from them or a slower rollout. But eventually the goal is like, hey, everybody would get this feature and clean up the feature flag and the branching in the code is gone versus in the configuration model is like, hey, those are actually going to change your business logic. And one example is what we were describing earlier, like, oh, you might want different granularity on when you're reconciling something and stuff like that. And that's just part of the like actual behavior in your product.
[00:37:42.23] - Robby
This episode of On Rails is brought to you by UnsafeSave, the fastest way to skip validations and write questionable data with confidence. Are you tired of pesky validations getting in the way of your momentum? Just add validate: false and boom, your record is saved. Empty fields, no problem. Missing associations, who cares? But wait, there's more. Act now and we'll throw in update_column totally free. Skip validations extensions and callbacks. That's right, bypass your business logic entirely. Update your database like no one's watching. Why wait for your code to behave when you can just skip the parts you don't care about? Unsafe save, because sometimes you just need that record in the database, consequences pending. When you think about how the application's been evolving over the last few years, how do you decide, uh, and this may be more directed for you, Eddie, when that customer flexibility can start to become for lack of a better term, technical debt of some sort?
[00:38:38.16] - Eddie
It's a good question. I think it's probably something that everyone struggles with. I think we do try to push back and think about the product in generic terms, right? Obviously, you don't want to be building one-off things. So I think it's more of putting yourself in their shoes and then also looking at it from our perspective, like, hey, is this something that makes sense in the product long-term and it's something that it's part of our process that you can standardize or do you actually work with the customer to be like, hey, like this is maybe another process that you can follow so we don't need to like add this feature and stuff. So I think it's kind of like both ways and like having empathy for like where they're coming from and like why they're asking for something. So with that, so like, I mean, I think we don't have all the answers. So sometimes they might ask for a feature and be like, oh, this actually makes sense. Like we need to add it. So I think it's probably not the best answer that everybody would want to hear, but I think it's really on a case-by-case basis, right?
[00:39:33.23] - Eddie
Like you just have to sit there and like think through the request and think about what makes sense for all of your customers and then where is your customer coming from.
[00:39:43.12] - Robby
No, that's true. I always think about how there's that balance of like, you're trying to, and I don't know if you're like a, it's like a sales-driven conversations like, hey, we can land this new customer and they're going to break, you know, like Yeah, we were 90% of what they think they need, but there's these few new things. So we're going to, we've committed to building these things like, well, how is that going to impact the rest of our customers we already have and will they need it or not? And how do you roll that out? And like everyone knows, I'm assuming it becomes its own, uh, interesting conference set of conversations and things to think about. But do you have like a pretty smooth workflow for like, is it, do you come in and like, all right, we're going to build out a new, is that, is that first of all, is that assumption true? That you need to do that kind of might be help bring in a new customer and you're like, we have to kind of change just a little bit for this customer. Is that true?
[00:40:31.04] - Eddie
Yeah, so I think fortunately for most of it, it's actually like a conversation with, in some cases with engineering and product, like actually jumping in with the customer and try to understand what they're asking for. So we were able to kind of work through the features a little bit more and not just essentially be like, oh, We promise we're trying to solve this, so you have to, you have to do it, uh, kind of thing.
[00:40:55.16] - Kagen
Yeah.
[00:40:55.18] - Robby
So you're part of the, you get to, the engineers get to be part of those conversations before that even, like typically before the ink even gets signed and you're like, this is what we can do and we'll make some adjustments to make this work for you or if necessary, or like, oh, we don't need that right now, but this is coming down the road. This is on the roadmap for the next 12 months and we'll get there.
[00:41:15.09] - Eddie
Uh, yeah. Yeah. So a combination of, uh, uh, of both those. So we. We do try as much as we can to like actually have engineers in the calls and so that you're close to the customer and you can actually understand what they're asking for.
[00:41:30.12] - Robby
Out of curiosity for, you know, just because I think a lot of people that work in certain types of SaaS businesses that don't work with like enterprise customers where you're maybe if you're selling $30, $50, $100 a month subscription service, there's like maybe some tiers, presumably you're not doing a lot of high-end customization for specific customers, because the idea is you're trying to get bulk customers, not necessarily larger type clients. Did either of you work at SaaS platforms before working at Ascend?
[00:41:59.08] - Eddie
I haven't.
[00:41:59.19] - Robby
Yep.
[00:42:00.15] - Kagen
Most, most of my experience is at different kinds of SaaS platforms. Yes.
[00:42:04.06] - Robby
And was it very, was it similar where there was a lot of this kind of, uh, per customer customization in those other environments as well?
[00:42:11.21] - Kagen
Yes. In particular for the largest contracts, we would be especially willing to, to do more custom things. Although I think something that holds for those companies as much as it does here is that generally before we go about agreeing to do something custom, we try to think about if that's going to generalize well to use cases for anybody else in the future, whether this makes sense as part of the product in general. And if it doesn't, we do generally try to push back on like super bespoke custom features that won't ever apply to anyone else. There are exceptions.
[00:42:49.06] - Robby
Sure.
[00:42:49.12] - Eddie
Sure.
[00:42:50.09] - Robby
Um, money is, money is an attractive thing to do businesses at times. Uh, depends if they'll pay for it, right? Uh, or pay, pay for some, pay towards that, I suppose. And there, I can appreciate that, but yeah, I was just always curious, like how that, what that actually looks like in different, in different environments. How, how do you think Rails simplifies that? Do you feel like this is more of a Do you feel like Rails is giving you enough tools to help work on these types of platforms like that where you are customizing it in that sort of fashion? Or do you think Rails should think about that a little bit more from the framework itself to have some more baked-in tooling of that sort?
[00:43:28.14] - Kagen
I think maybe contrary to the expected criticism that Rails's convention over configuration makes it more difficult to do something custom, I think that's more true on the tech side than it is on the product side. And because of the ability to spin up something new fairly quickly, relying on those conventions, it actually makes it a little bit easier to build out something bespoke pretty quickly as well as, and the duck typing also goes a long way in that respect too.
[00:44:01.23] - Robby
Can you give me an example of where duck typing might come into play in that, in that environment for our listeners that might not be that familiar with the term? Yeah.
[00:44:09.10] - Kagen
The duck typing helps with taking something that we have already built out, for example, that's shared across a lot of customers. And then we want to do something weird with it that is maybe oddly specific to a particular customer, but we can rely on a lot of the logic that's already been written because with duck typing, what duck typing essentially says is that the type of the object is its interface. It's what, what methods does it respond to? Rather than statically declared type information. And so if we're able to write some kind of adapter or interface to our already existing code that responds correctly to the pre-existing interface, we can leverage a lot of the work that's already been done on past projects.
[00:44:58.10] - Robby
Is that like as much as just like overriding a method to do something that— and you can kind of rely on everything else that's already there?
[00:45:05.23] - Kagen
Yeah, overriding a method or writing a class that matches the interface of some other module that we already have that does most of the work that we're looking for and expects that interface, for example. It really helps with hacky things like that in particular.
[00:45:26.18] - Robby
It sounds like you have a positive favor. I've heard some people say this is a can feel a little too magical or a little weird to kind of diagnose. And if, if for people coming in and they don't understand how these things work, yep. What's your take on that?
[00:45:41.06] - Kagen
That, that's a real trade-off. Um, especially if the module that you're writing is— you can definitely write it in a way that makes it harder to figure out what's going on if you're— especially if you're not following Rails conventions like the conventions exist for a reason. They help you find where things are, reason about where things work when you follow them. If you don't, some of the magic, so to speak, can kind of take over and make it a little bit more difficult to reason about what's going on. But that's just one of the trade-offs that you make when Ruby on Rails is fully leaning into the design philosophy that it believes in, which is like duck-typed OO.
[00:46:22.03] - Robby
Any— you have any thoughts on that as well, or? Um, nothing to add. Okay. Well, let's shift gears a little bit. I want to kind of pivot over to talk about infrastructure because you mentioned that you're running on Heroku and I think maybe for listeners out there, maybe an obvious question might be is like, or do you think there's a good chance you're going to try to stick around as Heroku as long as they provide support? Or I know there's a little bit of ambiguity about their long-term support right now that I've heard from, we've been hearing from their blog posts.
[00:46:49.14] - Eddie
I think we've been as confused as everyone else in terms of like, I think I saw it on Reddit of like, hey, you release a blog post saying you're going into maintenance mode, but you keep releasing features. So people were confused. So I think we're kind of in the same camp. I think we start exploring options and I think we'll move if we have to. I mean, in some areas we're kind of hitting the limits, but we still have to do the homework of like actually exploring other options. We use Autoscale for auto-scaling of our Rails application. And I remember those guys said that, oh, we're going to look into all the vendors and kind of write up a blog post to suggest something that's like Heroku-like. So I don't know if that has been posted already, but we're kind of waiting on that. But I think we'll also do all homework when the time comes.
[00:47:40.03] - Robby
So when we had prepped for this conversation, you mentioned that you needed to do some tuning with Puma. Can you speak to that specifically within the Heroku environment?
[00:47:48.10] - Eddie
Yeah, I think we actually took the time to probably take advantage of all the concurrency that you can do in Puma. So we were running into issues where we're kind of taking in requests and just like sitting idle and not doing anything in the process. So we started to take a look into that a little bit more and actually leverage what Puma does for you, which is like, hey, you can handle requests more concurrently and scale a little bit better. And that Turned out to be a pretty easy fix. It was just a configuration change. Like, hey, we increase that concurrency and all of a sudden you can have a little bit better performance with the trade-off of, I think we also had to make some memory modifications because then that would obviously start consuming more memory in your, in this case, in our Heroku dynos. So we also make a change there to start using, I don't know if I'm pronouncing it correctly, but jemalloc for Ruby memory management. So yeah, I think that those two have been working good for us so far, and the change really helped us increase our request per second processing.
[00:48:58.05] - Robby
Did you, how did you go about testing that out? Did you just tweak it a little bit and watch it, or did you, were you able to kind of simulate this a little bit in like a different environment before you rolled it out to the production stack?
[00:49:10.20] - Eddie
Yeah, so we were able to simulate it a little bit. So we, we have like a sandbox environment in Heroku that's like very configured pretty much like similarly to our production application. So we were able to test some of these changes. I think we've made the mistake in the early days of like, I think we actually tried to turn down a while back and I think the application just stopped working. So we learned our lesson.
[00:49:37.15] - Robby
That happens. You learn, you figure it out. Sometimes it works. Quite often it works. Are you relatively up to date with versions of Ruby and Rails and Puma?
[00:49:47.13] - Eddie
Yeah, I think actually, thanks to the fact that we haven't deviated a lot from like just your standard Rails application, I think the upgrades have been pretty easy and we were able to like keep up to date. I think now with all the coding agents and stuff, it's like pretty easy to like, hey, can you bump to this version, then have it run tests, then we run our own verification just for our own sanity and From there, it's like they're pretty easy to, to get through.
[00:50:13.10] - Robby
And out of curiosity, is your team kind of standardized on AI tooling at this point that you're using, or you're kind of an experimental— everybody's doing a little bit of trying different things at the moment?
[00:50:24.16] - Eddie
Um, I think more of the trying different things at the, at the moment. I would say the majority are kind of centered around, uh, Claude, uh, code. But yeah, I think that's still pretty open for the team, and you're also not enforcing like, hey, you have to use this or that.
[00:50:38.19] - Robby
And for our listeners and watchers, we're recording this in the second or the later half of June. So by the time this gets published, who knows where that, where things might be. These things are changing quickly. So I'm kind of curious, you were talking about your deployment workflows there. And you're using Heroku. So it sounds like fairly kind of straightforward deployment and you got sandbox environments. Have you built any interesting tooling internally to help you simulate realistic data in your local development environment, or what's that workflow look like so you can have that sort of volume? Do you— are you able to even test locally or in your sandbox with like a pretty big volume of transactions where you bring in— you mentioned like an account or a customer might have a transaction that were 10,000 transactions in a single, uh, I forget the word you're using, a supplier statement.
[00:51:30.10] - Kagen
Yeah.
[00:51:31.00] - Robby
So do you have tooling locally for that that you've been building out for yourselves or a bunch of Rake tasks or what's this look like?
[00:51:38.11] - Kagen
Yeah, we have seeding Rake tasks, but we haven't, to my knowledge, Eddie, correct me if I'm wrong, we haven't invested tremendously in a lot of guardrail-y type AI-specific technology such as like self-verification stuff. Um, or some of the things that you'll, uh, that some other organizations are working on to make it so that AI can fully develop and verify its own work. We haven't gone fully in on that. However, um, I think on a personal level, I've used it a lot for that purpose. A good example of exactly just what you were talking about, a little while ago I needed to migrate one of our CSV builders over to streaming from like building a CSV file in memory because the scale that we were hitting was causing some of the reports that we were— CSV reports we were having customers build. They can click a button essentially to export a CSV of any list view that they're looking at. And the volume was starting to get to the point where building it in memory would become insanely slow or sometimes even crash. And so I needed to migrate it over to streaming. I went ahead and implemented that and needed to figure out how to test it.
[00:52:56.16] - Kagen
And so it was just like, Claude, go into— open a Developer Rails console and just instantiate a statement with a million transactions. And then we'll generate a CSV to see if the memory profile of the generator is stable over time and that we got like a reasonable decrease in the amount of latency that it takes to process.
[00:53:17.08] - Robby
Hmm, that's kind of a cool use case. What about, are you able to test things out in like your sandbox environment to kind of before you push things out to production for customers when you're rolling things out? Like how much does that mirror production data if you're allowed to share?
[00:53:36.12] - Eddie
Yeah, so all of the data there is just mock data, but we actually rely a lot more heavily on the tests that we write, especially on our Rails application. It's not so much on like seeding, data and clicking around, primarily because the Rails application serves as a JSON API today. And I think when we're testing the boundaries more of a data scale, I think we do those more as one-offs. So, we have some scripts, so benchmark this, and I think just measuring memory pressure in different places. But yeah, I think we rely a lot more heavily on, hey, like making sure that we, that we follow a test-first approach. So that's kind of what prevents us more from the, from the bugs rather than like, oh, let's go into a sandbox and test something out.
[00:54:26.23] - Robby
What are you using on the, the view layer then? Are you, you mentioned there's a JSON API. Are customers using the JSON or is that something you're using internally with a different front-end framework or something? Or are you using Rails views at all?
[00:54:39.16] - Eddie
So we use it in Both internally and externally too, because we also have an API that customers can use. So we use Active Model Serializers for that side of the view layer.
[00:54:51.22] - Robby
And what about when it comes to rendering out HTML and CSS?
[00:54:55.08] - Eddie
So that one, we actually run a Next.js application.
[00:54:58.19] - Robby
Okay, so you have Next.js involved as well. Is that for the customer-facing side of things?
[00:55:05.14] - Eddie
Yes, they have a dashboard where they do most of their tasks.
[00:55:10.23] - Robby
Since that part of the application, what was the rationale for using that versus using what kind of comes with Rails at the time? Do you remember that?
[00:55:20.16] - Eddie
I think a lot of that was actually very also, the beginning from the engineer that was kind of setting the front end side of things. It was more of the ease of like, I think at that point in time, I don't remember if Hotwire was already out or not when we were starting that. But I think a lot of the dynamic work on the UI was a little bit more challenging to do, especially I think just like finding engineers in the market that were like, oh, I think that the Rails story at that point on the UI was a little bit more involved. So this was easier to spin up and get something done.
[00:56:00.18] - Robby
It's historically been been a little complicated over the years. Uh, like what's the new hot thing and is this gonna stick around? And, uh, what's the job description?
[00:56:08.20] - Eddie
Yeah, it has gotten a little bit, uh, uh, better.
[00:56:12.09] - Robby
Um, it's, it's true. It's, it's, um, but you're always kind of like wondering, I'm like, well, maybe, maybe I have enough people using this. Well, hopefully. Uh, yeah. You know, another thing you mentioned, um, when we were talking is like you mentioned like Sidekiq. Could you recall approximately how many How many jobs or Sidekiq jobs has your system been processing to date, approximately?
[00:56:34.21] - Eddie
Yeah, I think we're more than 1 billion. I think we're around like 1.3, 1.4 billion. So we, I think Sidekiq is a really important part of our infrastructure. We rely heavily on it, and I think we've liked it a lot.
[00:56:51.03] - Kagen
That's great.
[00:56:51.13] - Robby
And I'm assuming you're doing that on a pro account.
[00:56:54.07] - Eddie
I don't remember. I think we're actually on the Enterprise account.
[00:56:58.03] - Robby
Enterprise.
[00:56:58.18] - Eddie
Yep.
[00:56:59.01] - Robby
Sorry. Sorry. That's awesome. When you're thinking about, you mentioned like using gems and stuff like that in the community earlier. Does your team kind of have an ethos or approach about how you vet potential gems you might bring in? And do you feel like that's different? Do you feel like your thoughts on that has changed now? We're talking about the AI era now where that becomes a potential risk of bringing in some external dependencies now. Like, so historically, what has been your ethos about when do you, when you decide to build something yourselves versus bring something into the, into the repository, into your project and throw it in your Gemfile? And then two, is your thoughts on that changing at all at the moment?
[00:57:41.14] - Eddie
I think at the very beginning, we were following more of the approach of like, hey, whatever we bring in is still like, code that we own, even if it's not like on our code base. So you have to understand that you cannot bring a gem that's like super complex for maybe a tiny little problem that you're trying to solve. So I think that was a big piece of like, hey, it's still code that we have to maintain. Like whenever we have to do an upgrade, like we have to understand what's going on. If there is a bug there, like we have to be able to figure out what's happening. So I think that has been a big piece of it. But I think with that said, we don't shy away from, especially for things that are already implemented, right? Like we use, for example, PaperTrail for our audit logs on models. Like, hey, we know it's something that has been around for a while, is stable, and we feel comfortable using it. So I think it's probably a combination of that, of like, hey, we're going to own it. And then like, hey, is it something that other people are using is stable and you feel comfortable adopting.
[00:58:45.11] - Robby
Have you needed to navigate in these past 5 years many gem migrations where you, maybe the support was discontinued and you needed to migrate to allow you to take advantage of some new functionality with, from the, from those gems, or maybe it was a blocker for an upgrading Ruby on Rails or Ruby itself?
[00:59:04.06] - Eddie
I think we've been lucky that every time we try to do something, like there's already a PR open on the repository. So it has been mostly like, hey, we need to cut our own version until that gets a release. So I think that has been probably the most common situation that we've encountered rather than like, oh, this is not supporting like, I don't know, the latest Rails version or whatnot.
[00:59:23.18] - Robby
Okay. That tracks. Do you have any forked gems right now that you've needed, you're maintaining? You don't have to disclose which ones they are, but.
[00:59:31.18] - Eddie
I think, uh, I think now we've been able to cut it down to maybe just 2.
[00:59:35.23] - Robby
Okay.
[00:59:36.11] - Eddie
Yeah. I don't, I actually don't remember the number of how many, how many we depend on today.
[00:59:41.00] - Robby
Speaking to that, the point about bringing in dependencies, whether that be on the, even sounds like you're even maybe on the npm side, like how do you and your team think about bringing in dependencies? Do you feel like you're, are you having active conversations about that right now? About like, hey, like this is, this is a way that we could be potentially at risk? Not that you weren't before, but it feels like things could sneak in a little bit quicker under our nose.
[01:00:08.09] - Eddie
You mean from the existing ones that we use or from new ones?
[01:00:10.18] - Robby
Potentially from new ones.
[01:00:12.21] - Eddie
I think it doesn't happen that often because we actually don't add new gems that much at this point in time. I would say the PgBouncer example that we were walking through, that was probably the most recent one. Aside from that, I would say the problems throughout the application are very similar. I think we've been able to constrain ourselves to a handful of gems. I think with that said, with the ones that we have, it is a constant thing that's on our mind of like, oh, how do we make sure that those ones don't get compromised? I think that's something that's top of mind, especially on the field that we're in.
[01:00:50.00] - Robby
I appreciate that. Kagen, from your side of the world, I'm thinking about infrastructure and decisions that are kind of affecting your day-to-day. One of the things in our conversations, prior conversations we were talking ahead of this, was that you mentioned that you're needing to also maintain Elasticsearch, I think, was part of the equation with your Postgres. What are you using Elasticsearch for and what sort of problems was it solving? And are there any new challenges that that has since introduced as well?
[01:01:18.18] - Kagen
Yeah, that's right. We, we maintain, uh, two primary data layers. So we have Postgres and Elasticsearch. The reason we have Elasticsearch is previously alluded to. Many customers, they have millions of policies. Supplier statements can come in with like tens of thousands of entries on them. And most of the fields that exist, especially on a transaction, our users at some point will want to be able to filter or sort by that value. Elasticsearch is really good for that on large datasets, especially because when it comes to supplier statements, for example, the number of those in production is just going to continue to grow unbounded. So we rely on Postgres for more detail view rendering, and Elasticsearch renders our list views and provides filtering, sortability, search.
[01:02:15.02] - Robby
Are there any interesting challenges with that? Or like, if you have a pretty growing database, dataset, how much of that data is getting indexed and how do you keep that in sync?
[01:02:26.23] - Kagen
Yeah, both of those are big challenges. The first one that's an evergreen ongoing challenge is when you're trying to maintain the same data in both Postgres and Elasticsearch, you need to make sure that you keep it in sync. We use Active Model Serializer to serialize the documents into ES.
[01:02:48.18] - Robby
Okay.
[01:02:50.12] - Kagen
And a couple of complications there. If some of the fields rely on other models in order to render their values, one, we need to make sure that when those dependent models get updated, the parent's document gets reindexed so that the list view and the filters and sorting is all accurate. And the other is, I'm sure you've seen lots of active model serializers that are not very performant, um, but N+1 queries with relations and such like that. Um, and when those things happen, it can put a lot of strain on the database. It can slow the queues down. So we have to be very careful.
[01:03:26.22] - Robby
What would, what would be the impact of it not being totally in sync right away? Like what sort of, is this like a customer-facing issue or is it just, it'll eventually get updated? It's just like a UI primarily issue?
[01:03:39.16] - Kagen
It's, it's definitely can be user-facing, um, if severe enough, because Elasticsearch is used to render the list views. The worst case scenario is if something doesn't get updated, the list view will continue to render an outdated value, or if it's a newly created record, it just won't show it at all. And then when you click into the detail view for that, it'll show the updated information. And so the UI will disagree on the state of the world in two different places. This can happen if, for example, the queues get backed up for processing Elasticsearch indexing. It can happen if we forget to re-index a document somewhere that's important, like on one of those dependent models, for example.
[01:04:26.07] - Robby
Do you have any ways to like convey that to the user if there is— things are a little out of sync, or is it just something— how would you— how would they— how would you know or they know unless they raise a support request or something, be like, hey, I noticed that the dollar amount or whatever in the— when I did a search was this and it's showing a different number here. And that, I would imagine, if that's— there's a little bit of a versus like they just— the data changed and they just happened to search like immediately at the same time and they didn't. Is that pretty obvious or like how would you know that that's happening?
[01:04:58.12] - Kagen
Yeah, um, I think whenever there is a failure in this regard, it can be visible to, um, end users pretty quickly. They'll notice it, they'll report it to us and we will know. Um, because in some contexts, for example, we're extracting a new supplier statement and all of its entries, there's an expectation of eventual consistency. And so there's a little bit of a grace period that we're allowed before like items populating into the list view, um, will show up. We'll generally try to block UI elements when we're waiting on something to reindex. Flows. What's that?
[01:05:38.16] - Robby
What's that workflow look like? How do you— is it just like it knows that it's going to be rebuilt and so it just waits?
[01:05:44.17] - Eddie
Or—
[01:05:45.16] - Kagen
yeah, in the case of a supplier statement, when we're extracting— extraction, for the record, is a PDF document comes in and we use AI to parse the document and extract all of the individual entries that are on that statement. And create new models in our database for each one. There's a record for each transaction, etc. Um, and while that's completing, we mark a status on the supplier statement that says that it's currently extracting or matching. Um, and we have— we use Sidekiq batching to orchestrate all of the things that happen as a result of extraction. So matching, for example, will happen automatically. After extraction, we'll try to match each of the entries to the policies that they belong to. And after all of the Sidekiq jobs have completed, the Sidekiq batch callback will fire and it'll set— it'll unblock the supplier statement for user viewing.
[01:06:44.16] - Robby
Okay. Are you using any sort of like state machine approach to processing all that, or?
[01:06:50.11] - Kagen
Yeah, we do. For extraction in particular, there's state machine transitions. We use What's the module called? It's AASM for short. I forget.
[01:07:01.03] - Robby
Acts As State Machine
[01:07:02.16] - Kagen
Yes.
[01:07:04.04] - Robby
How does that kind of match up from your perspective, Eddie, when you think about product requirements and things around architectural decisions? Was there— or do you already have patterns like that that you're already using elsewhere in the, in the codebase? And or is these kind of more newer things that you've been adopting as a team?
[01:07:21.17] - Eddie
Yes, I think Those patterns are pretty like center across our application, especially because we do depend heavily on Elasticsearch. It is something that we keep an eye on so that the UI is actually reflecting what the customer is expecting to see. Otherwise it creates a bunch of confusion on their side and on our side as well. So yeah, I think those patterns are just like shared across the application.
[01:07:45.11] - Robby
Yeah.
[01:07:46.07] - Kagen
We have some patterns that help us enforce that. like callbacks on our models, for example, onSave or onUpdate that will run any related ES serializers. But it's still tricky when the document depends on other models for some of its values to make sure that when those change, any necessary documents that depend on it are also updated.
[01:08:07.09] - Robby
That's, that, that's challenging. Out of curiosity, have you looked at alternatives to Elasticsearch since you decided to first use it? I'm not saying I know something off the top of my head, but this seems like an interesting challenge for that other teams might be facing. With. And, um, so if anyone's listening out there and got ideas and want to respond or write something in the comments there, but is this something you're actively reviewing, or was just like, no, this, this would probably be fine for now? But I'm just thinking, you've mentioned hockey stick growth data consumption, so I'm like, that sounds like a pretty big index to be managing.
[01:08:39.22] - Kagen
We have, yes. Um, recently we started looking at ParadeDB, which is a Postgres extension We haven't made any decision to migrate or anything like that. It's a broader organizational decision, but it's something we started looking at because it's a Postgres extension. And so you get the indexing, the document re-indexing for free as part of just saving records to Postgres in the first place. Something interesting about it is that it's also relied on by, um, Modern Treasury, who do similar money accounting workflows, and they rely on it for like Elasticsearch type work as well.
[01:09:21.02] - Eddie
Interesting.
[01:09:21.19] - Robby
I'm not familiar with that one. I'll have to check that out. You always learn something new. Um, I'm kind of curious, you know, one of the other things I wanted to kind of come back to, you mentioned a little bit earlier around using AI and Kagen, you specifically had mentioned like finding like you're, you're able to, you're using AI a bit more. How would you describe it from within the realm of using it specifically working on Ruby in particular, like when it comes to using these LLM tooling?
[01:09:50.21] - Kagen
Yeah, I think there's a consensus opinion developing that AI models, LLMs tend to work best in the context of strongly typed languages because it provides more guardrails, um, and automatic verification for the AI to check its work. But I think on the other side of that, the conventions, the expressiveness of Ruby and Rails also lend pretty well to AI usage because they provide a ready-made set of rules for LLMs that kind of keep them on the rails, that are also token inexpensive. As anybody that has worked a lot with LLM coding knows, they are not very proactive about developing abstractions upfront or thinking about how things should scale into the long term or setting conventions. They will do exactly the thing that you ask them to do to solve the problem at hand. Rails and Ruby's conventions and expressiveness help keep the token expenditure down so that the reasoning stays better and also provide a good set of conventions that are ready-made for it to rely on.
[01:11:07.20] - Robby
We've definitely been hearing that more as of late. I'm looking forward to seeing some more studies on this so we can say this with some confidence, like we know what we're talking about. And I know there's been some useful anecdotes being shared out there. So I want to take that with a grain of salt for anyone listening. This isn't a scientific perspective, but something we're trying to read the tea leaves as best we can here. And Eddie, what's your kind of take on— what's your perspective here? How are like AI tooling changing like your workflows over at Ascend?
[01:11:39.01] - Eddie
I think it helps, uh, uh, like being a little more efficient with the, the task. And I think for actually one great output from all of this is I think people are kind of spending more time on expressing their thoughts of like, oh, I want to do, uh, this, and This is why, and this is what I'm thinking. So I think that part has been great because you have more of that clarity up front of like, hey, this is what I'm trying to do. On the other side, I've been following more of the approach and I think my opinion more on like, hey, whatever output comes from AI, like I still consider it of the output of the peer that I'm working with. Like I think I'm kind of against the like, oh, I'll just put some code there and like pray that somebody else will review it for me and stuff. I think that part I'm still like mindful of, of like, hey, whatever output is coming out, like I don't care if you wrote it yourself or if you use a tool to write it. Like I still expect for the person writing it to like understand it and know what's going on and all of that.
[01:12:39.06] - Eddie
Like I think I kind of see it the same way as like back in the day when you would like copy something from Stack Overflow. Like you still have to understand it. You can just like copy paste and like, oh, I'm done and move on.
[01:12:51.00] - Kagen
There's a blog from Ashby about AI usage actually that really resonates with me on philosophy about AI usage in software development in general. It's from this month and it's called AI, Ashby Engineering, and the Future. And the, the way that it positions it, the thesis is that the cost of writing code is heading towards zero. But the cost of producing meaningful software is not. Um, and that the differentiators for good engineers, which are judgment, taste, customer understanding, those things are becoming more important rather than less important. And we have to think more and think harder than we did before because LLMs make it very easy to avoid the cognitive burden of doing the work. That's the, um, some of the appeal sometimes is being able to offload the, the thinking part. But when LLMs are producing very plausible but not quite correct, or not quite the right pattern code, it requires us to think harder and better than we did before. And, uh, I really appreciated that framing.
[01:14:03.18] - Robby
Any follow-ups on that, Eddie?
[01:14:06.06] - Eddie
No, I think it, uh, it's gonna align with, uh, my, my thinking as well.
[01:14:10.21] - Robby
I think it's an interesting time right now, and I— for, for teams, and I've, you know, interviewed a number of people in some other upcoming episodes that'll publish in the near future, or I guess actually before this episode. But, uh, so just, it was always like trying to track like where, where different teams were at. And like you talk to some teams and they're like fully like, oh, we're just letting the, the, uh, as Kent Beck is now calling them, the genie is just, just, you know, he's, this is not what he, he wouldn't say this, but like I've talked to some people who are like, AI is cranking out the code and they're barely even reviewing the code anymore. And I'm just like, that's wild. And I'm like, how, what kind of guardrails do you have in place? Do you feel so confident about that? And then I'm like, it's just interesting. Like, and some people are like, we gotta be super careful. I'm not saying that's you, but it's just like, uh, where, where is, what's, what's, and maybe is everybody's reality reality and that's okay. And like, maybe it's always been, uh, we're always, we're always working a little bit differently.
[01:15:07.07] - Robby
I don't know. I'm, I'm, I'm trying to be optimistic and skeptical at the same time. And it's an interesting place to be.
[01:15:12.15] - Kagen
And I think there's a lot of nuance in it that's not always acknowledged by either side having the conversation. It's very context dependent. If the cost of being wrong is low, then it makes a lot more sense to just fully vibe it out and have the AI ship the code and barely review it. Because if, if it's wrong, then you can just iterate and ship again. But if the cost of failure is high and you need to get it right the first time, It makes a lot less sense to operate that way. And the article that I mentioned, the Ashby engineering one, also frames the AI usage on a spectrum. Full vibe coding for low-stakes tasks, like a throwaway script or some kind of internal tool. And we do this at Ascend. We use Lovable for internal tooling because the cost of making a mistake is fairly low.
[01:16:05.20] - Robby
Yeah, yeah.
[01:16:07.11] - Kagen
But on the other side, when the stakes are high, I think human-driven with like AI as more of a sidekick makes more sense. You can have AI type out some of the code, but you can review it very carefully, think through the problem, make sure you're using the right abstraction, et cetera. In our context, a lot of times the stakes are high because we're, we're dealing with money. Our customers are accountants. They care. About the details. And so even within the same company, you see the different spectrum of what's acceptable.
[01:16:40.16] - Robby
I think that's a good place to end the AI topic. I appreciate that. I'm gonna— I'll find out, find that link and share that in the show notes for us. For, as you said, AI Ashby, or I'll have you send that over to me and we'll include that for everybody in the show notes. Kept you guys long enough, so I want to kind of start working Towards kind of a few last things I want to, I'm kind of curious about, and this may be for you, Eddie. You know, you talked earlier about coming back to Ruby on Rails as part of the secret sauce of Ascend's success to date. And it's, it's a scent, uh, so to speak. Uh, and what other technical decisions do you feel like your team's made that you feel like is really glad you did aside from Ruby on Rails, obviously, is there something else you feel like you, you, you were able to do early on that you felt like has been a really good anchor for the organization on the engineering side of things.
[01:17:34.01] - Eddie
I mean, I think tied to kind of the Rails stuff, I think on the infrastructure side, we've also tried to sort of keep things minimal. I think from the very beginning, our main focus was like, hey, how can we spend more time thinking about the product and building stuff? So we try to minimize the extra stuff as much as as possible. And it doesn't mean like, oh, we're not going to pay attention to it, but it's just more like we'll pay attention to it when it matters. So I think we operate a lot under that principle of like, it is not engineering for the sake of engineering. I think it's just more like if we have a product and a customer base to serve, so like we want to dedicate our time to that. So I think a lot of the infrastructure decisions come with that of like, hey, you have the trade-off of like, oh, so it might be more expensive than what you pay if you roll out everything on your own. But at the end of the day, like, hey, we don't have a dedicated like infra or DevOps person. So we have to keep things simple so that we can keep moving forward.
[01:18:43.09] - Eddie
So I think we operate a lot under that principle. And I think also on enforcing standards. I think that was something that we tried to do at the beginning. We want things to keep a cohesive approach and stuff. And I think with that also came the testing philosophy of, hey, we're writing things. We want to start writing tests for everything that we do, especially on our backend when we're dealing with money and all of that. So I think that kind of ingrained with the team and it's something that we keep up until this day. And I think something that has saved us also a lot from like, I think just having bugs in production and stuff. I think just allows you to move with velocity at the same time.
[01:19:34.05] - Robby
That's great. Out of curiosity, I don't think you touched on this earlier. What testing framework are you using within your Rails app?
[01:19:40.16] - Eddie
Today we use RSpec.
[01:19:42.13] - Robby
Okay, using RSpec.
[01:19:44.02] - Eddie
All right. Yeah, I think that decision is probably debatable, but okay, we have it now, so we're sticking with it.
[01:19:49.14] - Robby
Uh, well, I've heard some people have been using, um, their LLMs to, to migrate, uh, and putting that like, we'll just offload some of that stuff. I don't know how— I guess if you can find some good, uh, parity there, that would be not the bad— not a bad thing necessarily. But, um, I'm a long time RSpec user and fan myself, but I've been using Minitest more recently on some some smaller projects and be like, I could, I could do this. I can probably get behind this as well. Um, but we don't need to have a, uh, a religious war about tabs or spaces or RSpec and, uh, Minitests there. But, um, no, that's useful. And I think, you know, your, your point there around maybe intentionally, uh, not having certain types of roles, maybe around like infrastructure and stuff like that and using something like Heroku and trying to use that as an interesting constraint for your organization. Maybe there's a tendency to kind of like, well, we'll add more complexity to the infrastructure if we have people that know how to do that. And maybe there's a really good reason for it, but then there's the cost of not maintaining that.
[01:20:48.21] - Robby
And people like, and maybe people on your team may or may not have the time and bandwidth to learn all those things. So also be able to participate in some of the infrastructure management that might go and get involved in that. Is that, is that a safe kind of read on what you were saying there?
[01:21:02.22] - Eddie
Uh, yeah, no, I think that's, uh, that's right. I think. I feel like all of the things that we do need a reason of like, oh, why are we doing it? But some of the problems, like for example, on like the data ingestion and Elasticsearch is like, hey, we have a reason because we're like now getting more and more customers. So like now it makes sense for us to spend time there versus like before, if we had like very little traffic, like why would you have like a super complex set of infrastructure when it's not actually necessary?
[01:21:29.03] - Robby
As someone that works in the consulting space where we get brought into projects. There have been many, many projects where, uh, there'll be 3 developers that had worked on an app and there'll be 5 repositories, Terraform involved, and I'm like, you've got no customers, what are you doing? Like, so that, that's a thing. Um, all right, well, with that, a couple quick last questions. Uh, we'll start with you, Kagen. Is— for both of you though, is there a programming book or technical book that you find yourself recommending to peers?
[01:22:00.15] - Kagen
Oh man, tough question. I think I need a second to think about one. Do you read programming books?
[01:22:09.06] - Eddie
I do.
[01:22:10.00] - Kagen
Not at the velocity that I wish that I did, but I do read them. One of my favorites is the classic, I'm forgetting actually the title of it embarrassingly, but it's an introductory programming book through the perspective of Scheme. From an MIT course. I don't know if that rings a bell for you, but that's a personal favorite of mine. Eddie, do you know the title of that one?
[01:22:34.08] - Eddie
I know, but I don't know the title.
[01:22:36.13] - Robby
What about you, Eddie? Do you have a—
[01:22:38.13] - Eddie
I read the thing a few years back already. It's A Philosophy of Software Design. It's a pretty, like, small book, but it has a lot of interesting pieces. One of those that you could probably read in a day or even in a week. It doesn't have a lot, but I think it just has like ideas around like how to design software that are pretty, I find pretty valuable.
[01:23:01.12] - Kagen
Structure and Interpretation of Computer Programs. That's the one. Okay. It's a, it's a real nerd's book.
[01:23:08.05] - Robby
Okay. Real nerd's book.
[01:23:09.05] - Kagen
All right.
[01:23:09.10] - Robby
Well, great. We'll definitely include links to both of those in the show notes. And out of curiosity, where does Ascend have like an engineering blog or anything like that at this point? Or have you talked about that at all? To share what you're cooking up over there?
[01:23:24.22] - Eddie
We've talked about it. We've never done it, but yeah, it's probably something that we'll look into.
[01:23:30.14] - Kagen
Yeah.
[01:23:30.15] - Robby
Maybe you can have one of your AI bots build you one. And then when your first post could be like, hey, we were on the On Rails podcast and check out the video. Here's some of the things we talked about. So I'm giving you your first blog post right there. But I think, yeah, I think there's something you folks are doing some really interesting stuff there that I think the community would benefit from hearing more about. I want to thank you both for taking time to come on On Rails and talk shop with us and share some things that are going on behind the scenes over at Ascend. I really appreciate that. And on behalf of our audience, thank you.
[01:24:03.03] - Eddie
Thanks for having us.
[01:24:03.18] - Kagen
Thanks for having me on.
[01:24:06.18] - Robby
That's it for this episode of On Rails. This podcast is produced by the Rails Foundation with support from its core and contributing members. If you enjoyed the ride, leave a quick review on Apple Podcasts Spotify, or YouTube. It helps more folks find the show. Again, I'm Robby Russell. Thanks for riding along. See you next time.
People on this episode
Podcasts we love
Check out these other fine podcasts recommended by us, not an algorithm.
Maintainable
Robby Russell
Remote Ruby
Chris Oliver, Andrew Mason, David Hill
The Ruby on Rails Podcast
David Hill
IndieRails
Jess Brown & Jeremy Smith