Cloud computing is the predominant infrastructure for new development, but bore and more organizations are adopting a hybrid approach to their IT environment. Hybrid implies the combination of multiple types of resources or technologies, typically involving a mix of on-premises infrastructure, private cloud, and public cloud services. It allows organizations to leverage the benefits of both on-premises and cloud-based solutions to meet their specific requirements.
This approach has a built-in appeal for larger organizations that rely on mainframe systems because of the rich legacy of mission-critical applications that run on mainframes. Mainframes are designed to handle large-scale, high-performance, and mission-critical workloads and excel in processing and managing vast amounts of data and transactions. So, organizations that have invested in mainframes rely on the platform’s high reliability, availability, and security.
At the same time, cloud computing offers tremendous benefits for new development of web and mobile applications on a cost-effective platform. The cloud is well-suited for many applications, particularly those that benefit from scalability, flexibility, cost-efficiency, and accessibility.
This means that organizations are keeping and extending their mainframe applications while also building out new applications using cloud services. Such an approach is called a hybrid environment or sometimes hybrid cloud computing.
Mainframes Generate Lots of Useful Data
Mainframe data plays a crucial role in a hybrid cloud environment, especially in organizations that have legacy systems and modern cloud-based infrastructure.
The importance of data cannot be understated because it is foundational for organizations to make informed decisions, optimize processes, and deliver better user experiences. For example, data is at the core of AI and Machine Learning (ML) development as models are trained on data sets to learn patterns and make predictions. And that is only one example.
Data enables businesses to gain valuable insights through business intelligence (BI) and analytics tools. By analyzing customer behavior, market trends, and operational metrics, businesses can make data-driven decisions to improve efficiency, optimize processes, and identify new opportunities. Data-driven marketing strategies rely on customer data to target the right audience with relevant messages. Data can be used for predictive maintenance by analyzing equipment performance information in real time to predict when machinery is likely to fail and proactively schedule maintenance. Data is used for fraud detection and risk management, personalizing customer experiences, and, indeed, to run day-to-day business operations. Truly, data is a requirement as every business and application uses and produces data.
Given its rich heritage, the mainframe is a phenomenal source of valuable data for all types of modern development. But it can be challenging to access mainframe data out of context. Consider the numerous different types and formats of mainframe data that exist.
Data may be stored on many different Database Management Systems (DBMS) on the mainframe. Even though Db2 for z/OS is the leading mainframe DBMS today, several other popular DBMSes are used to power mainframe applications and store critical data, including IMS, IDMS, Adabas, and Datacom/DB. These DBMSes all store data differently and use multiple different models, including relational, network (where relationships between records are defined through sets and pointers), hierarchical (where data is organized in a tree-like structure with parent-child relationships), and others. And let’s not forget that many organizations run Linux for z which means that many popular Linux DBMS products – such as Oracle, MySQL, PostgreSQL and others – run on mainframe in Linux partitions.
But mainframe data need not be stored in a DBMS, and indeed, much of it is not. The predominant form of data structure on the mainframe for years (even before the advent of database systems) was flat files, also known as QSAM or physical sequential files. Flat files contain records with no structured relationships and require additional knowledge to interpret their content (for example, a COBOL copybook with a file description including fields, data types, and lengths).
Another prevalent type of mainframe data is VSAM or Virtual Sequential Access Method. VSAM data can be accessed using an index or sequentially from direct access storage devices. There are three ways to access data in a VSAM file: random (or direct), sequential, and skip-sequential. As with flat files, VSAM files require a file definition to access, as there is no embedded description of the data other than perhaps a key.
Another type of mainframe data that may be useful exists in log files. Log data is used with DBMSes to manage and record changing data, but it can also be used by transaction processing systems (such as CICS and IMS/TM) and other system software. The operating system, z/OS, also writes log data to multiple locations, such as the SYSLOG, the job log, the OPERLOG, the console, and more. All of these logs are formatted differently and can be difficult to interpret without additional context and documentation. Nevertheless, log data can be a valuable tool for system management and uncovering useful operational information.
We must also acknowledge that not all mainframe data will be stored on disk. Mainframes often utilize magnetic tape storage for archival purposes. Tape data only can be accessed sequentially. Of course, much of the “tape” storage used on mainframes these days is virtual, meaning that it is on disk, but mimics traditional tape processing.
Obviously, the mainframe is a rich source of data. But how can we access this data in a hybrid environment from applications that may not be running on the mainframe?
Accessing Mainframe Data
There are many different ways for hybrid applications to utilize mainframe data. The key is enabling the application to understand and access the data in a way that makes the most sense for its operations.
One approach is to use Application Programming Interfaces (APIs) or Web Services to interact with mainframe data. These APIs enable authorized access to specific mainframe functions, data repositories, or transactions. Cloud applications can use standard protocols like RESTful APIs, SOAP (Simple Object Access Protocol), or WebSphere MQ to communicate with mainframe systems and exchange data.
Another popular mechanism is deploying Middleware and Integration Platforms as intermediaries between cloud applications and mainframe systems. They provide connectors, adapters, or APIs specifically designed to interface with mainframe environments. These platforms facilitate seamless data integration, transformation, and messaging between cloud applications and mainframe data sources.
Message queues and event streams also allow cloud applications to access mainframe data. MQ technologies like IBM MQ or Apache Kafka can be used to exchange messages or events with mainframe systems. Mainframe applications can publish messages to the queue or subscribe to event streams, allowing cloud applications to consume and process the data in near real-time.
Or cloud applications can connect directly to mainframe databases using the appropriate database connectivity protocols. For example, Db2 on z/OS supports industry-standard database connectivity options like ODBC (Open Database Connectivity), JDBC (Java Database Connectivity), and SQLJ (SQL in Java) for accessing mainframe data. Cloud applications can use these interfaces to query, update, or manipulate mainframe databases directly.
Another possibility is to move the data from the mainframe to the cloud using ETL (Extract, Transform, Load). ETL procedures can be deployed to extract data from mainframe systems, transform it into a suitable format, and load it into cloud-based data storage or analytics platforms. This approach involves extracting mainframe data through various methods such as file transfers, database queries, or APIs. The data is then transformed and loaded into cloud data warehouses, data lakes, or other storage solutions for further processing and analysis. An alternate, similar approach is ELT (Extract, Load, Transform), in which data is ingested before it is transformed.
Yet another approach is to use Replication and Synchronization mechanisms to continuously capture changes made to mainframe data and propagate them to corresponding cloud data repositories in near real-time. Examples of this technology include IBM’s Data Replication or Change Data Capture (CDC) and Precisely’s Connect product. This approach enables cloud applications to work with up-to-date mainframe data without directly accessing the systems.
Finally, Virtualization and Emulation techniques may be employed to create an abstraction layer between cloud applications and mainframe systems. This involves running mainframe operating systems or applications on virtualized environments or emulators within the cloud infrastructure. Cloud applications can use standard networking protocols or APIs to interact with the emulated mainframe environment.
The chosen method depends on factors such as the nature of the mainframe data, security considerations, performance requirements, existing mainframe infrastructure, latency considerations, and compatibility with cloud technologies.
The Bottom Line
It is possible to combine several of these techniques to integrate cloud applications with mainframe data, enabling modernization, data access, and seamless interoperability between mainframe and cloud environments. But one thing is for certain… you cannot ignore the mainframe as you implement your hybrid cloud environment!
