Business Continuity Policy
1. INTRODUCTION
1.1. The activities related to the business of an institution are subject to recent interruptions for the most diverse reasons, and it is necessary to prepare for an adequate reaction to these issues. This is an integral part of corporate ICT governance which is structured on the principles and best practices of internal and external service management.
1.2. The business continuity plan consists of a set of action plans aimed at providing a basis for obtaining confidence in the organization's business in a consistent and recognized manner in accordance with its business continuity management capacity.
2. OBJECTIVES AND BENEFITS
2.1. The objective of the Business Continuity Plan is the formalization of strategies capable of carrying out recovery, continuity and resumption in times of crisis, preventing the institution's critical processes from being affected by generating impacts, resulting in organizational resilience. The benefits of an effective business continuity plan include:
2.1.1. Ability to proactively identify impacts of an operational disruption.
2.1.2. Efficient response to interruptions, resulting in minimizing the impact on the organization.
2.1.3. Maintenance of the ability to manage risks that cannot be insured.
2.1.4. Encouragement of work between teams.
2.1.5. Ability to present a possible answer through a testing process.
2.1.6. Improving the reputation of the institution
3. BUSINESS CONTINUITY MANAGEMENT (BCM) MACRO-PROCESS
3.1. This topic addresses the macro-process of ATTRUS' Business Continuity Plan, and should be used as a normative/guiding instrument for the execution/creation of procedures and actions to carry out recovery and resumption in times of crisis that may occur within the scope of its activities.
3.2. The methodology used for the construction of the Continuity Plan used materials based on ANSI/ASIS ISO 22313:2020, guide for the use of ANSI/ASIS ISO 22301, which aims to provide recommendations on good business continuity management practices. The basis of the recommendations is the business continuity life cycle, shown in the figure below:
3.3. The management of the program enables the establishment and maintenance of the business continuity capacity in an appropriate manner, so the active participation of senior management is essential for the BCM program to be properly introduced, established and supported as the culture of the organization.
3.3.1. Its management is separated into three stages:
- Attribution of Responsibilities: Which consists of selecting those responsible for the BCM program in each unit and their respective responsibility in the process;
- Implementation of business continuity: In this stage, the program is disseminated to stakeholders, training for the team, and continuity tests are carried out. The actual implementation should use a project management method already adopted by the organization;
- Ongoing management: The following activities constitute this step: Define the scope, roles, and responsibilities of BCM; Appoint an appropriate person or team to manage ongoing BCM capacity; Maintain the current business continuity program through good practices; Promote business continuity throughout the organization broadly, where appropriate; Administer the testing program; Coordinate critical analysis and regular updating of business continuity capability, including critically analyzing or redoing risk assessments and business impact analysis (BIA); Maintain documentation appropriate to the size and complexity of the organization; Monitor the performance of the business continuity capacity; Establish and monitor change management and management succession regime.
4. UNDERSTANDING THE ORGANIZATION
4.1. Understanding the organization, identifying its core products and services and critical activities and the respective resources that support them. This ensures the compliance of the Continuity Plan with the organization's objectives, obligations, and legal responsibilities.
4.2. It is necessary to fully understand the interdependence between the activities that the organization performs, as well as external relations with other organizations. But for this to be done, it is necessary:
4.2.1. Identify the organization's objectives, stakeholder obligations, legal duties, and the environment in which the organization operates.
4.2.2. Identify the activities, assets and resources, including external ones, that support the delivery of these products and services.
4.2.3. Assess the impact and consequences on the failure time of these activities, assets and resources.
4.2.4. Categorize your activities according to your recovery priorities.
4.2.5. Identify and assess threats that may disrupt key products and services and the assets, activities, and resources that support them.
4.3. Once the institution's internal understanding is completed, it is appropriate to carry out an assessment of the impacts on:
4.3.1. People's well-being;
4.3.2. Damage to or loss of facilities, technologies or information;
4.3.3. Failure to comply with duties or regulations;
4.3.4. Reputational damage;
4.3.5. Damage to financial viability;
4.3.6. Deterioration in the quality of products or services;
4.3.7. Environmental damage.
4.4. It is convenient for the organization to estimate the necessary resources during the recovery of each activity, whether it is people, dependencies, technology, information and supply (external services and suppliers).
4.5. Not all risks can be identified and mitigated, or even if they are, the ability to perform an action ahead of it may be limited or its cost may be disproportionate to the potential benefit.
4.6. In its analysis, the hypothesis of acceptance of this risk, transfer of risk (contracting insurance) or even the change, suspension or termination of the related service, product, activity, function or process must be considered.
5. DETERMINING THE BUSINESS CONTINUITY STRATEGY
5.1. This step continues the "Understanding the organization" element, because using the result of the analysis previously carried out, it is possible to evaluate a set of strategies so that the most appropriate response for each product or service is chosen.
5.2. The strategies chosen should focus on the resilience and countermeasures already in place in the organization. Aiming at its continuity considering an acceptable level of operation and its respective time.
5.3. For each resource of the organization, the following must be outlined: maximum tolerable interruption period of the critical activity; Implementation costs of the selected strategies; and Consequences of not acting.
5.4. The organization's resources must have identified and appropriate strategies, as consolidated in the following table:
| RESOURCE | RESPECTIVE RESOURCE STRATEGIES MAY INCLUDE |
|---|---|
| People | documentation of the method of execution of critical activities; multidisciplinary training of employees and service providers; separation of key skills to reduce risk concentration (this may cause a physical separation of employees with key skills or ensure that more than one person has them); use of third parties; succession planning; e retention and knowledge management. |
| Facilities | alternative facilities (environments) within the organization, including the relocation of other activities; alternative environments provided by other organizations (through or without reciprocal agreements); alternative environments provided by specialized third parties; work from home or remote locations; other locations that are agreed upon as appropriate; e use of alternative workforce in an established location. |
| Technology | geographic distribution of the technology, i.e., keeping the technology in different locations that will not be affected by the same business disruption; storing older equipment as a replacement in case of emergencies; e additional risk mitigation for single equipment or for a long lead time. Specifically strategies related to IT services: target recovery time (RTO) of systems and applications that support the fundamental activities identified in the BIA; location and distance between technological facilities; number of technological installations; remote access; use of empty facilities (unstaffed) instead of occupied facilities; telecom connectivity and redundant routing; nature of the "fail over" (whether manual intervention is required to activate the alternative Tl features or whether this should occur automatically); and third-party connectivity and external links. |
| Information | Any and all information in physical (printed) or virtual (electronic) form that is directly linked to the performance of critical activities in the organization must have an adequate level of: confidentiality; integrity; availability; e update. |
| Supplies | increase in the number of suppliers; recommending or requiring suppliers to have a validated business continuity capability; contractual obligations and/or service level agreements with key suppliers; or the identification of alternative suppliers that are able to meet the demand. |
6. DEVELOPING AND IMPLEMENTING A BCM RESPONSE
6.1. This stage is directly related to the development and implementation of the preparations previously made so that in the event of any incident there is a minimum structure that allows the organization to:
- Confirm the nature and extent of the incident;
- Take control of the situation;
- Control the incident;
- Communicate with stakeholders.
7. TESTING, MAINTAINING, AND CRITICALLY ANALYZING BCM PREPARATIONS
7.1. At this stage, tests and critical analyses of the previously implemented plans are carried out. Testing is of paramount importance for the development of teamwork, competence, confidence and knowledge, vital factors in the occurrence of an incident.
7.2. It is appropriate that the tests can advance a predicted result that has been planned in advance and included in the scope, in addition to allowing the use of innovation in the organization.
7.3. The main objective of conducting tests is to seek the best assurance that the Business Continuity Plan works as expected when necessary, therefore, they must:
- To test the technical, logistical, administrative, procedural and other systems in operation aspects of the PCN;
- Test BCM preparations and infrastructure, including roles, responsibilities, and any incident management locations and work areas, among others; e
- Validate the recovery of technology and telecommunications, including the availability and redeployment of personnel.
7.4. Among the advantages of improvements in the Business Continuity Management capacity, the following can be listed:
1. Exercise the organization's ability to recover from an incident;
2. Verify that all critical activities of the organization, its dependencies and priorities are covered by the PCN;
3. Highlight premises that should be questioned;
4. To build trust in the participants involved in the test;
5. Increase awareness of the business continuity process by the organization through the publication of the test;
6. Validate the functionality and timeliness of the process of restoring critical activities; e
7. Demonstrate the competence of the titular incident response teams and their surrogates.
7.5. It is worth noting that all tests must be realistic, carefully planned, aligned with stakeholders. Your objectives, scales, and complexities should be clearly defined, generating post-test reports and analysis. As shown in the following table:
| COMPLEXITY | TEST | PROCESS | VARIATIONS | RECOMMENDED FREQUENCY |
|---|---|---|---|---|
| Simple | Table test | Critical analysis/correction | Update/Audit/Verification | Annually |
| Simple | Go over the steps of the plan | Questioning the content of the PCN | Include interaction and validate roles | Semiannually |
| Medium | Test critical activities | Execution in environment controlled | Execution of operations selected from an alternate location | Semiannually |
| Medium | Simulation | Use hypothetical situation to validate that the plan has all the necessary information | Incorporation of associated plans | Annually |
| Complex | Test the entire PCN | Testing involving all areas | Annually |
7.6. The maintenance of BCM consists of critical analysis questioning any assumptions adopted for any component of management. As well as its wide dissemination updated, corrected or amended. The analysis can be carried out by internal or external audits, or even through self-assessment, focusing on verifying the points listed below:
7.6.1. All key products and services and the critical activities and resources that support them have been identified and included in the organization's BCM strategy;
7.6.2. The organization's BCM policy, strategies, structure, and plans accurately reflect its priorities and requirements (the organization's objectives);
7.6.3. The organization's BCM competence and capacity are effective and adequate and will enable the management, command, control, and coordination of an incident;
7.6.4. The organization's BCM solutions are effective, up-to-date and adequate, as well as appropriate to the level of risk faced by the organization;
7.6.5. The organization's BCM maintenance and testing programs have been effectively implemented;
7.6.6. BCM strategies and plans incorporate improvements identified during incidents and testing and in the maintenance program;
7.6.7. The organization has an ongoing BCM training and awareness program;
7.6.8. BCM procedures have been effectively communicated to the relevant team and whether this team understands their roles and responsibilities;
7.8.9. Change control processes are in place and work effectively.
7.7. An NCM self-assessment process plays an important role in ensuring that the organisation has robust, effective and adequate competence and capacity. This process qualitatively verifies the organization's ability to recover from an incident. A self-assessment should be carried out that verifies the objectives of the organisation and takes into account the relevant standards and good practices.
8. BUSINESS CONTINUITY PLAN
8.1. Concepts and definitions
8.1.1. Information assets: means the means of storage, transmission and processing, information systems, as well as the places where these means are located and the people who have access to them.
8.1.2. Business continuity: the company's ability, both tactical and strategic, to plan and respond to incidents and business interruptions, in order to maintain its operations at an acceptable level, previously defined, minimizing their impacts and recovering losses of information assets from critical activities.
8.1.3. Data center: all space in which the Information Assets are located, as well as their auxiliary structures such as UPS, Battery Bank and Electric Power Generator.
8.1.4. Disaster: a sudden, unplanned event that causes loss to all or part of the organization and has serious impacts on its ability to deliver essential or critical services for a period longer than the target recovery time.
8.1.5. Incident: event that has caused any damage, put at risk, any critical information asset or interrupted the execution of some critical activity for a period shorter than the objective recovery time.
8.2. Presentation
8.2.1. The Business Continuity Plan (BCP) ensures the continuity of its business in the event of a stoppage resulting from an accident in one or more processes considered critical, and must establish scenarios of unexpected situations or incidents, whether operational, disasters or crises. The continuity plan will act as a response to the results of the Business Impact Analysis and Risk Analysis, having the duty to manage them, giving due attention to:
- Strategic, tactical, and operational alternatives to respond to disruption;
- Prevention of new losses or unavailability of priority activities;
- Details on how and under what circumstances the company will communicate with stakeholders.
8.3. Purpose
8.3.1. The aim of this Business Continuity Plan (BCP) is to promote effective and fast protection strategies and measures for critical IT processes, in order to ensure their preservation after the occurrence of a disaster, until resumption in a timely manner. The BCP will act as a response to the results of the Business Impact Analysis and Risk Analysis, providing which actions will be carried out at each stage of the plan.
8.3.2. The process of elaboration of the PCN was based on standards and research from other institutions, within the best practices for the construction of a Business Continuity Management System. The activities were developed by employees and managers, in a collaborative way, with face-to-face meetings to align the plan with institutional policies.
8.3.3. This plan is divided into 4 (four) other stages, which are:
1. Crisis Management Plan (PAC) - Defines the roles and responsibilities of the teams involved with the activation of contingency actions, before, during and after the occurrence;
2. Contingency Plan (CP) - Defines the most immediate needs and actions. It should be used only when all preventions have failed;
3. Disaster Recovery Plan (DRP) - Determines the planning so that, once the contingency is controlled and the crisis is over, the original levels of operation are resumed and;
4. Operational Continuity Plan (OCP) - Its objective is to reestablish the operation of the main assets that support the institution's operations, reducing the downtime and the impacts caused by an eventual incident.
8.3.4. Among the objectives of the Plan, the following procedures stand out:
8.3.4.1. Identify all IT business processes, defining critical activities and classify them;
8.3.4.2. Identify and document risks that may compromise the continuity of critical activities;
8.3.4.3. Identify threats, vulnerabilities and estimate risks;
8.3.4.4. Identify existing controls;
8.3.4.5. Identify, document and assess possible impacts on the continuity of critical activities, if such risks materialize;
8.3.4.6. Determine and calculate the time and cost of business downtime and recovery;
8.3.4.7. Define, implement and maintain a formal and documented process for the Business Impact Analysis;
8.3.4.8. Assessment of the impacts of not carrying out critical activities over time;
8.3.4.9. Setting deadlines in a prioritized manner for the resumption of activities, at a minimum tolerable level of execution, taking into account the time in which the impacts of the interruption become unacceptable;
8.3.4.10. Identification of interdependencies and resources that support the activities, including suppliers, third parties and other relevant stakeholders;
8.3.4.11. Determine appropriate business continuity strategies to protect, stabilize, continue, resume and recover priority activities, as well as their interdependencies and support resources (Crisis Management Plan);
8.3.4.12. Establish appropriate levels of authority and competence, in order to ensure effective communication to stakeholders, as well as to ensure the continuity of critical activities (Business Continuity Plan);
8.3.4.13. Enable the continuity and recovery of critical activities in the event of an interruption (Disaster Recovery and Contingency Plan);
8.3.4.14. Conduct training and evaluations of the PCN periodically to ensure the maintenance and proper functioning of the continuity plans;
8.3.4.15. Conduct tests to ensure the efficiency of business continuity;
8.3.4.16. Promote employee awareness;
8.3.4.17. Identify opportunities to improve business continuity.
8.4. Model Plan (PDCA)
8.4.1. The plans defined herein will follow the "PLAN-DO-CHECK-ACT" (PDCA) Model to plan, establish, implement, operate, monitor, critically analyze, maintain and continuously improve the effectiveness of the System.
8.4.2. The PDCA model will help in the continuous improvement of the Business Continuity Plan:
- Plan - Follow a business continuity policy, objectives, goals, controls, processes, and procedures relevant to the improvement of business continuity, in order to have results aligned with the objectives.
- Do (implement and operate) – Implement and operate the business continuity policy, controls, processes, and procedures.
- Check - Monitor and critically analyze performance in relation to business continuity objectives and policy, report results to management for critical analysis, define and authorize improvement and correction actions.
- ACT (Maintain and Improve) - Maintain and improve the BCP by taking corrective and preventive actions, based on the results of the management review and re-evaluating the business continuity scope, policies and objectives.
8.4.3. For each of the stages, Action Plans must be made, and these must be prepared as soon as they occur, based on their temporality and impact. These should form a log or record of actions, so that for each event it is possible to verify what was done at other similar times.
8.5. Validity of the PCN
8.5.1. This plan will be valid for four (4) years.
8.6. Revisions
8.6.1. The review of the plan will be carried out in the following situations:
8.6.1.1. A maximum of 2 (two) years;
8.6.1.2. At times when ATTRUS deems it necessary;
8.6.1.3. Depending on the results of the tests carried out; or
8.6.1.4. After the occurrence of any event or significant change in the information assets, activities or any of its components.
8.7. Invoking the plan
8.7.1. This plan will be activated when a disaster occurs, in the event of an unknown risk or if a vulnerability is highly likely to be exploited. The plan may also be activated when the need for tests occurs or by determination of ATTRUS.
8.8. Main risks
8.8.1. The BCP was designed to be activated when there is an occurrence of disasters that present risks to business continuity or essential services. Below is the table that defines these risks, as well as points out the parameters to report the possible causes of the occurrences.
| DISASTER EVENT | POSSIBLE CAUSES |
|---|---|
| 1 - Power interruption | Caused by an external factor to the electrical network of the building or its location with an interruption duration of more than 12 hours. Caused by an internal factor that compromises the building's electrical network with short circuit, fire and infiltrations. |
| 2 - Unavailability of network/circuits | Fiber optic rupture resulting from the execution of public works, disasters or accidents. |
| 3 - Human error | Any act caused by negligence, recklessness and/or malpractice. |
| 4-Insider attacks (disgruntled employees) | Attack on Data Center Assets or Internal Servers |
| 5 - Fire | --------- |
| 6- Natural disasters | --------- |
| 7- Hardware failures | Failure that requires replacement of part or whose repair or acquisition depends on budget. |
| 8-Cyber attack | A cyberattack that compromises the performance, data, or configuration of critical services. |
8.9. Critical processes and systems
8.9.1. Critical processes and systems can be defined as a work process that, once paralyzed for a longer time than defined by the business managers, will significantly affect operations, generating an impact on customers. This impact is defined by the following formula: BAT = RTO + WRT.
- MTD (Maximum Tolerable Downtime): Defines the total amount of time that a business process can be interrupted without causing any unacceptable consequences. This value must be set by the Disaster Committee. Different business functions will have different MTD's.
- Recovery Time Objective (RTO): Determines the maximum tolerable amount of time required to bring all critical systems back online (e.g., restore backup data or fix a failure).
- WRT (Work Recovery): Determines the tolerable amount of time required to verify the system and/or data integrity (verify databases and logs, for example). When all systems affected by the disaster are scanned and/or recovered, the environment is ready to resume production again.
| CRITICAL PROCESS | MTD | RTO | WRT |
|---|---|---|---|
| Hardware Defect | 4h | 3h | 1h |
| Damage, loss, or corruption of servers, computers, and operating systems | 4h | 3h | 1h |
| Backup failed | 10 am | 8 am | 2h |
| Human error (recklessness, negligence and/or malpractice) | 3h | 2h | 1h |
| Failure of internal equipment | 2h | 1h | 1h |
| Structural failure (physical damage to the building/office) | 2:30 am | 2h | 30min |
| Absenteeism of essential employees | 1h10 | 1h | 10min |
8.10. Business Impact Analysis (BIA)
8.10.1. A business impact analysis (BIA) predicts the consequences of disrupting a business function and processes and gathers information needed to develop recovery strategies.
8.10.2. Potential loss scenarios shall be identified during a risk assessment. Operations can also be disrupted by the failure of a supplier of goods or services.
8.10.3. BIA shall identify the operational and financial impacts resulting from disruption of business functions and processes. Possible scenarios and impacts to consider in the event of process/business disruption include:
- Sales and lost income;
- Contractual penalties;
- Customer dissatisfaction or defection;
- Physical damage to the building;
- Damage or breakdown of machinery, systems or equipment;
- Restricted access to a location or building;
- Shutdown of public services (e.g., power outage);
- Damage, loss or corruption of servers, computers, operating systems, applications and data;
- Absenteeism of essential employees.
8.11. Time and Duration of the interruption
8.11.1. The Business Impact Analysis exists to define parameters on the time required to recover the services (maximum acceptable unavailability/objective for recovery time) and the time required for their backups (objective for Recovery Point / maximum data loss). These parameters will be calculated after the preparation of questionnaires and tabulation of risk assessment.
8.11.2. The questionnaires are designed to obtain information for the preparation of:
- Critical business systems/processes under the company's responsibility;
- Raise the investment and funding for the implementation of alternatives to avoid interruption and/or recover operations
- Impacts to be considered;
- Degree of criticality of critical systems/processes;
- Target time and maximum downtime, as well as its positive point;
- Decide on alternatives, resources and their costs based on cost-benefit analysis;
8.12. Conducting the BIA
8.12.1. In order to correctly conduct the BIA, it is first necessary to analyze the criticality and then to evaluate all impacts (financial, legal, operational, administrative, image and human resources).
8.12.2. The table below serves as a reference to identify the threats, their impacts, what is the value on the business, their importance and what procedure will be adopted to correct this threat.
8.12.3. The revision of this table must be made annually, or whenever there are significant changes in the activities, in which case ATTRUS must request the revision.
| THREATS | IMPACT | VALUE | IMPORTANCE |
|---|---|---|---|
| Hardware Defect | Direct | High | 1 |
| Damage, loss, or corruption of servers, computers, and operating systems | Direct | High | 1 |
| Backup failed | Direct | High | 1 |
| Loss of business-critical data | Direct | High | 1 |
| Attacks on systems | Direct | High | 2 |
| Human error (recklessness, negligence and/or malpractice) | Direct | High | 2 |
| Software Outdated | Indirect | Medium | 3 |
| Internal equipment failure | Direct | Medium | 3 |
| Structural failure (building/office damage) | Indirect | Low | 4 |
| Absenteeism of essential employees | Direct | Low | 4 |
8.13. BIA Report
8.13.1. To verify the level of criticality of the risk, the following formula will be used: THREAT + IMPORTANCE = BUSINESS IMPACT
8.13.2. To calculate this formula, the following scoring parameters will be used:
Importance 1 - 15 points for each process;
Importance 2 - 10 points for each process;
Importance 3-07 points for each process;
Importance 4- 03 points for each process.
8.13.2. The result of this formula signals the degree of impact of the non-implementation of the contingency and/or its delay, in the event of a disaster or service stoppage.
| RESULT | SEVERITY | IMPACT OF NON-IMPLEMENTATION |
|---|---|---|
| 70 to 100 | Critical | High |
| 40 to 70 | Moderate | Medium |
| 10 to 40 | Lightweight | Low |
8.14. Crisis Management Plan - (CMP)
8.14.1. This plan specifies the actions in the face of disaster scenarios. Actions include administering, managing, eliminating or neutralizing the impacts inherent to the relationship between those involved and/or affected, until the crisis is overcome.
8.14.2. The objective of the PAC is to ensure communication, manage crises, and enable a linear understanding of the actions before, during, and after the occurrence of a disaster for all those involved. The specific objectives of the PAC are:
1. Ensure the safety of people's lives;
2. To guide employees and other collaborators on the conducts that will be taken;
3. Inform customers with clarifications consistent with what happened in a timely manner;
4. Minimize inconvenience on the unfolding of incidents and stimulate joint efforts to overcome the crisis.
8.14.3. In the event of a disaster, it will be necessary to contact the affected areas to inform them of its effect on the continuity of services and time for recovery. The plan should include actions to redirect incoming phone calls to a second number. The communication team will be responsible for contacting customers and other affected parties and passing on the relevant information. The communication will take place as follows:
8.14.3.1. Notify the authorities: The competent authorities must be notified in the event of a disaster that involves risk to people, providing information on the location, nature, magnitude and impact of the disaster.
8.14.3.2. Communicate the responsible sectors: In addition to communicating to those responsible, you must also inform:
- Nature, impact and scope of the disaster;
- Contingency actions in progress;
- Processes/systems and services covered by the continuity plan
8.14.3.3.Communicating suppliers/service providers
8.14.3.4. Communicating external collaborators
8.14.3.5. Notify all of the above parties when operations return to normal.
8.14.4. Once the return of the essential functions of the system and its total stability has been validated, as well as the stability of the datacenter, if this is the case, the Communication Team will contact all those involved described in this plan, providing the return information and the status of the essential services, and must issue an opinion reporting the activities carried out to restore the services.
8.15. Contingency Plan - (CP)
8.15.1. This plan aims to establish a recovery after a disaster, with the objective of ensuring the reestablishment of essential systems and their respective activities. Its main objective is to list the procedures defined to allow data processing and storage services to continue to operate, even with a certain degree of degradation.
8.15.2. The contingency plan is defined by three macro pillars, which are based on:
- People: deals with the human resources involved in contingency activities;
- Organizations: deals with the availability and security of organizational structural resources to support the necessary activities in contingency;
- Technology: deals with hardware and software resources supported by technologies and complements to meet the contingency.
8.15.3. Following this line, we have as a reference four groups, distributed as follows:
8.15.3.1. Contingency of physical infrastructures: Includes disaster situations, natural or not, such as floods, landslides, fires, failures in the energy supply, among others. In general terms, these are occurrences that prevent access and/or use of the facilities, as well as relevant physical damage to facilities and/or equipment, intentional or not.
8.15.3.2. Contingency of people: These are those where key employees are not present due to strikes, illnesses, leaves, etc.
8.15.3.4. Contingency of external services: Comprises situations of non-provision of contracted service considered critical to the processes.
8.15.4. In order for the contingency to follow its flow, it is recommended that these steps be present. They are:
8.15.4.1. Diagnosis: consists of identifying weaknesses that could be the focus of problems for the company's IT sector.
8.15.4.2. Risk analysis: based on vulnerabilities, possible threats and factors that may lead to the materialization of these risks must be considered, such as virus attacks and the absence of a corporate antivirus.
8.15.4.2. Definition of priorities: identify the company's vital processes and point out which systems need to be recovered first or preferably in case of problems.
8.15.4.3. Determination of strategies: this is the way to define how each system should be recovered (using software or applications), when and who are responsible for it.
8.15.5. The plan will be terminated as soon as all services are stable and the operation of essential systems operating normally. The team responsible for the return must issue an opinion reporting the activities carried out, which in turn must provide a communication of return to the activities.
8.16. Disaster Recovery Plan - (DRP)
8.16.1. This plan describes the scenarios of inoperability and their respective procedures, so that, once the priority activities are defined to reestablish the level of operation of the services, the contingency is controlled and the crisis is over, the organization returns to its normal levels of operation. To ensure the return of operations after the occurrence of a crisis or disaster, the objectives of the recovery plan are:
8.16.1.1. Assess damage to data center assets and connections and provide means for their recovery.
8.16.1.2. Avoid the unfolding of other incidents.
8.16.1.3. Reestablish the data center within the tolerable timeframe.
8.16.2. For the plan to take place as planned, the following steps must be performed:
8.16.2.1. The team responsible for backups and servers must identify and list all damaged assets from the occurrence of the disaster;
8.16.2.2. The network team must identify the interruptions of connections and access generated after the disaster, informing whether the coverage is in the local network, WAN or with the service provider;
8.16.2.3. Those responsible for the PRD must map which services have been discontinued containing the information on loss of asset and connection;
8.16.3. The committee responsible for the PRD, after mapping the losses and impacts, will prepare a recovery schedule for the investments, taking into account the following applications for recovery:
8.16.3.1. Replacement of assets and equipment: In the event of loss of assets, the need to acquire lost assets that cannot be recovered must be immediately informed. The team will measure how long the acquisition will impact each service, communicating if there is any alternative solution to be taken while the acquisition is being carried out.
8.16.3.2. Reconfiguration of assets and equipment: The responsible team must verify that the configurations of the repaired or replaced assets are in full operation. If they are not, provide an estimated schedule to configure these assets.
8.16.3.3. Environment Testing: The primary environment (site and/or data center) shall be tested prior to data recovery to ensure that the recovery process occurs as planned. Tests and recoveries should:
- Ensure the same levels of capacity and availability of essential services before the disaster;
- Ensure the integrity of data, which may be corrupted or outdated;
- Validate all previous settings;
- Support the return of systems according to demand;
- Verify data integrity and restore backups if necessary.
8.16.4. The plan will be terminated as soon as the recovery procedures are carried out by all teams. At the end of all procedures, the service recovery information will be consolidated in a specific opinion, informing the time of reestablishment of each service, equipment acquired and/or relocated, if applicable, suppliers that had to be activated recovery procedures carried out, among other relevant information.
8.17. Operational Continuity Plan - (OCP)
8.17.1. This plan describes the outage scenarios and their respective planned alternative procedures, defining the priority activities to ensure the continuity of services and restore the operation of the main assets that support IT operations, reducing the downtime and the impacts caused by an eventual disaster.
8.17.2. The purpose of the plan is to ensure that continuity actions during and after the occurrence of a crisis or disaster, in the case of contingency actions only, are intended to maintain the continuity of business processes and vital services. It is through this that the process teams will know how to act in the absence or failure of any component that supports it, thus ensuring the continuity of the process, reducing its impacts.
8.17.2.1. Provide means to maintain the operation of the main IT services and the continuity of operations and essential systems;
8.17.2.2. Establish alternative controls, rules and procedures that enable the continuity of IT operations during a crisis or disaster scenario;
8.17.2.3. Define the forms, checklist and reports to be delivered by the teams when executing the contingency.
8.17.3. Once the occurrence of an incident, crisis or disaster is identified, the operations and backup team must verify the dimension of the impact, extent and possible consequences of the event. After the disaster impact assessment, the responsible team must fill out a questionnaire for evaluation and decision on the activation of the plan and the start of contingency actions. This questionnaire should be disseminated to all teams involved.
8.17.4. Given the approval for the activation of the plan by those responsible, an emergency meeting will be called with the leaders in order to:
- Coordinate deadlines and orchestrate contingency actions;
- Inform the teams of contingency actions with the prioritization of essential services.
8.17.5. In order to resume the business, it will be necessary to verify the following steps:
1. Estimate the impact of data loss;
2. Identify affected assets;
3. Map assets to be recovered;
4. Estimate the volume of data to be retrieved;
5. Recovery time and possible operational losses;
6. Implement recovery procedure;
7. Test procedures performed;
8. Pass on the procedures to the servers and verify improvements.
8.17.6. The plan will be terminated as soon as the operation of the essential systems is validated, as well as the data center, if applicable, reporting its stability and normality. After this process, an opinion will be issued by the responsible team, informing what led to the activation of the plan, the activities carried out and the resources that were used, in order to then communicate the stability of the system to all sectors.