CREATING A WEB INTERFACE FOR MONITORING THE STATUS OF COMPUTING CLUSTERS
Abstract and keywords
Abstract:
The paper considers the development of a lightweight web interface for monitoring the status of computing clusters operating under the control of a specialized task queue management system. The relevance of the solution is due to the need for a simple and effective tool for operational control of distributed computing resources in a scientific and educational environment where heterogeneous and low-power systems are often used. The architecture of the solution based on the analysis of text files of statistics and their automatic transformation into HTML pages is presented. The system is based on the use of Bash and Python scripts, which ensures minimal dependencies and ease of deployment. The basic functionality of the system is described, including displaying the current task queue, monitoring the load of computing nodes (processors, memory), detecting failures based on timestamp analysis, and creating a history of completed tasks. Special attention is paid to timestamp processing methods to ensure correct calculation of task execution time and algorithms for cross-platform compatibility (Linux, Windows). The developed solution has been successfully tested in real-world operating conditions and has demonstrated stable performance over five years on several computing clusters. The results of testing the system are presented, confirming its high efficiency: the page generation time for a cluster of 10 nodes is less than 1 second, the error detection accuracy reaches 98%, and reliability over the period of operation is estimated at 99.8%. The key advantages of the system are low demands on computing resources, ease of deployment, adaptive design for mobile devices and support for multicluster architecture, which makes it especially attractive for educational institutions and small research groups.

Keywords:
WEB INTERFACE, MONITORING, COMPUTING CLUSTER, JOB SCHEDULING SYSTEM, PBS, PYTHON, HTML, DISTRIBUTED COMPUTING, HIGH-PERFORMANCE COMPUTING, ENERGY EFFICIENCY
Text
Text (PDF): Read Download
References

1. Gentzsch, W. “Sun Grid Engine: Towards Creating a Compute Power Grid.” Proceedings of the 1st International Symposium on Cluster Computing and the Grid, 2001, pp. 35–36.

2. Zhuravlev, S. S., Rudometov S. V., Okolnishnikov V. V., Shakirov S. R. Application of Model-Driven Design to the Development of Automated Process Control Systems for Hazardous Industrial Facilities. NSU Bulletin. Series: Information Technologies, 2018, Vol. 16, No. 4, pp. 56–67. DOIhttps://doi.org/10.25205/1818-7900-2018-16-4-56-67.

3. Altair Engineering PBS Professional 2022 User’s Guide. 2022, 284 p.

4. DeLeon, R. L., Furlani, T. R., Gallo S. M., et al. XDMoD: A Tool for Comprehensive System Monitoring and Performance Analysis of High-Performance Computing Centers. Proceedings of the 2015 Winter Simulation Conference, 2015, pp. 3233–3244. DOIhttps://doi.org/10.1109/WSC.2015.7408425.

5. Staples G. TORQUE Resource Manager. Proceedings of the 2006 ACM/IEEE Conference on Supercomputing, 2006, p. 8. DOIhttps://doi.org/10.1109/SC.2006.21.

6. Henderson R. L. Job Scheduling Under the Portable Batch System. Job Scheduling Strategies for Parallel Processing, 1995, vol. 949, pp. 1–6.

7. IBM Platform Computing. IBM Platform LSF. [Electronic resource]. URL: https://www.ibm.com/products/hpc-workloadmanagement (accessed October 15, 2023).

8. TORQUE Resource Manager. [Electronic resource]. URL: http://www.adaptivecomputing.com/products/torque/ (accessed October 15, 2023).

9. Jette M., Yoo A., Grondona M. SLURM: Simple Linux Utility for Resource Management. Proceedings of the 2002 Cluster World Conference, 2002.

10. Thain D., Tannenbaum T., Livny M. Distributed Computing in Practice: The Condor Experience. Concurrency and Computation: Practice and Experience, 2005, vol. 17, no. 2–4, pp. 323–356.

11. Massie M., Chun B., Culler D. The Ganglia Distributed Monitoring System: Design, Implementation, and Experience. Parallel Computing, 2004, vol. 30, no. 7, pp. 817–840.

12. Barth W. Nagios: System and Network Monitoring. 2nd ed. No Starch Press, 2008.

13. Zabbix LLC. Zabbix Documentation. [Online]. URL: https://www.zabbix.com/documentation/ (accessed October 15, 2023).

14. Grafana Labs. Grafana Documentation. [Online]. URL: https://grafana.com/docs/ (accessed October 15, 2023).

15. Elasticsearch B.V. Kibana Guide. [Electronic resource]. URL: https://www.elastic.co/guide/en/kibana/current/index.html (accessed October 15, 2023).

16. Masliy, A.A. “Creating a Web Interface for Monitoring the Status of Computing Clusters.” In: Engineering Thought. Proceedings of the II Republican Scientific and Practical Conference. Kazan, 2024. pp. 62–64.

Login or Create
* Forgot password?