Home / Finance / What if the system fails: The problem is not just yours, developers feel it is theirs too

What if the system fails: The problem is not just yours, developers feel it is theirs too

Today, there is hardly any business that does not rely on technology and software solutions. By this, we do not only mean products and services but also that the foundation of all businesses consists of software, various applications, and solutions that are, in fact, tools for performing daily tasks. A prime example is email, but today almost every segment of business is digitized and connected to some software and/or analytics. To better illustrate how to build a reliable software system in an unpredictable environment, we will present an example from the banking sector and our own experience.

Conditions for flawless operation

The core of the banking system is the foundation of all business activities of any banking institution. It contains the essence of the business, from client data to all transactions that are conducted. In other words, it is a system that centralizes all necessary information about clients, enabling the operation of products that are further personalized to enhance the user experience and through which all financial and operational activities such as risk management, financial assets, mobile and online payments pass.

Just like other services that banks offer today. This core system can – and must – communicate with other systems around it, such as ATMs, customer service, mobile applications for end-users, many internal records, and tools used by clients, as well as bank employees. Therefore, the core of the banking system is the heart of that financial institution and must quickly resolve all inquiries, whether internal or external.

This leads us to the conclusion that a reliable core, or a reliable software system, is essential for every successful bank. Because of such a core, the bank is functional. A reliable system actually enables the provision of the best user experience, and that is why the entire business is reliable and efficient. A reliable system is one that is simultaneously efficient, financially viable, and easily and quickly accessible. On the other hand, the system must be secure.

In the context of the financial sector, especially when third-party services such as online and digital payment tools are involved, it is a system that must function flawlessly, 24 hours a day, seven days a week. Just like online stores and various other applications that have become part of modern life, which actually have a software core in the background. In caring for such a system, it is crucial to maintain the infrastructure in the full sense of the word so that all its parts work flawlessly at all times. On the other hand, to maintain and improve such a system, it is necessary to ensure the appropriate staff.

Stable external partners

For anyone developing software or introducing new software into business, the key challenge is to create a reliable software system due to the challenges posed by unpredictable conditions in the operational part, which is daily exposed to new disruptions. Service interruptions are very difficult to prevent, although most risks can be assessed and prevented in time, but there are situations that no one can influence. When viewed from the perspective of the development team, especially an external one, it is they who must react quickly and resolve the problem in a very short time, enabling the system to function.

Teams of external partners engaged in maintaining and improving the core of the system are usually composed of people who are very reliable and have been in this business for a long time. This is necessary not only for client satisfaction but also for the overall business, primarily due to the stress that such situations cause employees, as well as financial losses when the business is at a standstill. That is why providers of such services have very stable teams with rich experience in software system development and extensive knowledge of specific business processes, as well as the industry.

In the event of an incident

Since, despite all knowledge and preparations, the operation of the core system sometimes fails, it is important to know how to behave in such a situation. Based on experience, we have learned several key things that could be useful to you regardless of the industry in which you operate. First of all, it is important to be aware that reliability does not stem only from architecture and application but also from knowledge and understanding of how to manage and react when failures occur, that is, when the system crashes.

It is also important to know how to set boundaries around the core (banking) system so that it can handle an incident. One of the key steps in quickly resolving an incident is monitoring and reporting problems. The established English terms are help desk and ticketing. This is a system that will allow thorough tracking of every request and timely reporting on the status, whether it concerns the software or hardware part of the system. If you have a good tool for this, you are already on the right path to setting up maintenance correctly. Various help desk tools are available, but in choosing, it is most important to ensure that the selected ones fit well into your processes and procedures.

For example, submitting a request for new functionality within the ticketing form is a very practical way to track requests. This way, you can monitor all requests in one place and have a clear overview of their statuses. To make this as effective as possible, we have introduced the practice that when submitting a request, or ticket, the client describes what it is about, providing as many details as possible. This is essential for good analysis and understanding of the root cause of the problem. Then, after the team takes on the task, we continuously update the status, or progress of the task, and when the problem is resolved, we can easily send a notification to everyone that the incident has been resolved.

In addition to the operational and communication aspects, it is important to continuously improve the reliability of the system; of course, to minimize incidents so that they are almost nonexistent. From the perspective of teams working on maintaining a system that is large and so complex, it is important to analyze every problem and learn from them.

Improving reliability

Based on this, active work should be done to ensure that there are no more incidents. That is why we are very focused on the quality of our service, striving to extend the uninterrupted operation of the system. We are extremely dedicated to studying all aspects of a problem when it arises, thoroughly examining its causes in a way that we have developed over the years, which has proven to be thorough. It is important to learn from mistakes. It would be better to learn from others’, but in any case, the most important thing is to strive to prevent problems.

Timely and clear

Unpredictable failures, or problems, are inevitable, especially in computer systems. Just remember Murphy’s Law: a problem will arise when you (or the client) need it the least. Here we come to another important thing: communication. And that is timely and transparent communication with all participants. Equally important is the mutual communication among team members as well as communication with the client. It is extremely important to establish communication between the client and the team, ensuring that the client is satisfied with it, enjoying it regardless of whether it is about daily work or problem-solving.

Such a development team understands the sense of urgency and the broader business picture. It must actually understand the client’s expectations very well. All team members must also be well connected and aligned. That is why technical knowledge, personality, and social skills are equally important in team composition. These skills particularly come to the fore when other departments are involved.

At that point, it is important to exchange information effectively so that everything flows smoothly and on time. Specifically, my team takes incidents very seriously. We have very well-structured communication and monitoring methods that we apply during incident resolution and in daily operations. We use all these methods to inform users about the status of service interruptions and share all the latest updates about the incident. Therefore, clients must feel that ‘their’ problem is ‘your’ problem.

Communication and collaboration with clients have helped us build mutual trust for lasting relationships and successful business. Additionally, we are committed to using proven practices that ensure high quality in software development and maintenance services. This ensures that every step, as well as the entire process, is optimized, efficient, and reliable. If I had to highlight the two most important things for building a reliable business software system, or those that most influence the stability of the system in unpredictable conditions, they would be communication and proven practices.

Tagged: