Evergen Azure to AWS Migration. Reducing cloud cost 20x times and increasing scalability 10000x times.
Evergen is a pioneer of smart energy market.
Evergen is one of the coolest companies you can meet on a market today.
What is special about that? Evergen is trying to revolutionise the energy market. What does it means?
Energy market is extremely complex and complicated topic, but even without the deep domain knowledge you defenetelly knows that the energy market is chanign right now: we are moving away from fossil fuel to reneewable sources: solar, wind, waves, hydrogen.. Australia is one of the leader in this green movement. And there is a reason for that:
And it looks like moving from centralised energy production to decentralised is requre a lot of changes in energy management as well.
This is why Evergen was created. The company was created 8 years ago to efficienly manage energy. Evergen is already much more than that, introducing virtual power plants, energy market trading, fleet management and more. But the core idea is still the same.
What is efficient energy managemnt? Let me talk thorugh three simple examples, they are very simplified, there a lot of more moving parts behind those optimisations. But we use simple examples to describe the core idea.
Imagine you have connected solar panel and bettery? What’s next. When you have a lot of sun - you don’t pay for electricity. Whe you have too much energy - you sell it back to power grid, reducing system costs.
But imagine that tomorrow is going to be a rainy day. We know that as forecasts are highly accurate, around 96-98%. That is why we can charge the battery at night with a low tarrifs and use the energy from battery during the day, to compensate low energy production from solar.
Another examples is about swimming pool. Swimming pools are popular in Australia, especially in the Northern regions. But the largest part of maintain is electricity bill. We can optimise it as well. Instead od running cleaning system constantly, we can switch it off for 15 mins every hour. Reducing energy costs 25% straight away. We can increase it ene more for chilly days, using temperature sensors, humidity sensors and weather forecast.
The third example is mind blowing. When you have exceed solar energy your energy retail obligated to buy it from you. But they buy it for minimal bargain price, 20-30 times!!! lower they sell it to you. And the reason for that is that you sell it in uncontrolled unpredicted way, in daytime, when the energy consumption is low. What Evergen does, it can predict energy prices. For example we know the Melbourne Cup is going to be tomorrow and everyone will flood the pubs, all the kitchens will make food and all TVs will be switched on. Or we could analyse weather forecast: hight heat or chill weather, when people will use heaters or air conditioners. Moreover getting detailed metrics from previous years we can predict seasonal or other consumption changes. Using this AI and Data model we can pre-charge the batteries for our customers and sell it when the electricity price will be on peak as an energy producer for market price. We have a real cases when people earn 40$ just in one day.
And all of those examples are simplified versions. There are much much more happening there. That is why I call Evergen one of the most innovative and interesting companies on an energy market.
The before. From proof to concept to real business scale.
All those magic, Evergen does, was not possible in the beginning of the company. First the company have to build Proof of concept. And the tech team build the solution on Azure using the tools they are familiar with with no intention for scale and cost optimisation.
The first Evergen’s architecture looked like that:
[[ The hight level pucture of Azure architecture ]]
As you can see there it is a monolith, and the monolith center is a databse. The majority of logic was implemented insite MSSQL stored procedurec. Which tells a lot about a power of MSSQL itself.
But there are disadvantages of that solution for sure.
Scalability: this setup culd scale vertically only. Yep, MSSQL can scale horizontally, but as any other relation datbaase it has some limitations. Talking about stored procedures, there is hard to scale.
Complexity: it was hard to debug, monitor and maintain those fucntions, As obviously it is misuse of the technology itself. You could not build in google analytics or tracing inside MSSQL stored procedure. It is extermelly powerfull tool, but it is not a regular backend environment as well.
Price: Talking about computing power of MSSQL everyone could egree that the costs of computation is hundreds times more expensive if we run the logic in a bare metal as a NodeJS and other funciton.
As you can see, the things woked well, for 200-300 active monitoring sites. But even for this amount the bill for the setup was 20-30k per month.
First step - extracting login to k8s.
The first step was extremelly obvous, we need to extract the logic from store procedures to serverless funcitons or microservices. We chosen the microservices, as we had a lot of new features we have to deliver, and having microservice architecture allowed us to work in parallel: improving the current infrastructure and delivering new features.
Microservices architecture requres orchestration. And teh obsous decision was to use K8S. Kubernetes is cloud agnostic, cloud-native, widely used and extremelly flexible.
Azure Managed Kubernetes service (AKS) was our choice. Is it the only option? No, there are some to mention:
Azure Kubernetes Services is the most complicated one, compare to alternatives. But it has some critical advantages:
First red flags
So the process started, we had a K8S cluster, it workes fine, we sarted to move some service from MSSQL to golang and nodejs service to K8S.
The process was painfull, as we had no documentation and no engineers who implemented the initial state. So we reverse negineered the stored procedured with all the dependencies and reimplement it, then run both in parralells and compare the outcome. If for some significant time (2-4 week)) outcoume is the same, it means the new service is works fine.
Sound easy, but we always hit the different result in the first iteration, half of the reasons was because of hidden dependency and hidden logic in stored procedures. The other half of the reasons that stored procedures had bugs. Nobody is doing proper integration testing for PoC solutions.
So we moved more services to k8s and at some point we hit soft limit on worker instances. It is a common practice to save unexpierenced cloud user from paying a millions for huge resources running in a cloud.
We contacted support team and decided that the next day the issue will be resolved.
But we were too optimistic.
Not next day, not even the day after, not even next week. We waited for TWO WEEKS. No we did not just sitted and waited, we contected support team every single day and got promises that they will solve that next day.
Could you imagine that? We begged every day to allow us to pay little more money to Azure. Unbeliveble. This is such a Microsoft style.
But 💩 happens. And Support team are people. And we could understand that. So this case was in the past and we moved next.
And could you imagine, when one day we come to the office and realised that all our k8s clusters are down! And the reason is that our cluster is outdated and Azure console is not supporting this version of kubectrl any more.
Have we received email about that before? Yep we had. It was one of hundred of emails we received regularly from Azure. It looked the same and it has no indication that the message is critical.
The problem was that we could not resolve the issue ourselves. To upgrade the cluster - we need running kubectrl. But Kubectrl was down because Azure console does not supported that version any more.
We spend 10 hours on a phone with Azure support, and you can say, that it would be easier to repeploy everything to the new cluster. But we had the issues with that as well: soft-limits - remember? We were not able to double our resources.
That was unexpected, unpredictable and uncontrolled situation.
This is was the point, when we started to talk about moving to other cloud loudly and seriously.
