Site Güvenilirliği Mühendisliği İçgörülerinden Yararlanarak Yerel Bulut İzleme Üzerine Bir Araştırma


Creative Commons License

Koç C., UĞURLU B., Karasulu B.

Muş Alparslan Üniversitesi Fen Bilimler Dergisi, cilt.14, sa.1, ss.29-37, 2026 (TRDizin)

Özet

Ölçeklenebilir ve taşınabilir uygulamalara yönelik artan talep, dağıtık sistemler ve bulut bilişimin önemini artırmıştır. Güvenilirlik ve erişilebilirliği sağlamak için Site Reliability Engineering (SRE) yaklaşımı kritik hale gelmiştir. Bu çalışmada, Amazon Web Services (AWS), Google Cloud Platform (GCP) ve Microsoft Azure’un yerel bulut izleme servisleri SRE çerçevesinde karşılaştırılmıştır. Golang tabanlı, mikroservis mimarili, API Gateway ve olay odaklı desenleri kullanan bir PaaS uygulaması geliştirilmiş; ilişkisel veritabanları ve bulut-yerel mesajlaşma servisleriyle desteklenmiştir. İzleme araçları, sistem sağlığını değerlendirmek üzere üçüncü taraf çözümlerle birlikte entegre edilmiştir. Analiz; çeşitlilik, maliyet, uyarı gecikme süresi, veri toplama sıklığı ve saklama süresi ölçütlerine odaklanmıştır. Bulgular, GCP’nin ayrıntılı sanal makine metrikleri sunarak ek yapılandırmaya ihtiyaç duymadığını ve olay odaklı mimarilerde en yüksek mesaj işleme kapasitesine sahip olduğunu göstermektedir. Çalışma, güvenilirlik odaklı dağıtımlar için platformların izleme yeteneklerine ilişkin karşılaştırmalı içgörüler sunmaktadır.
The demand for scalable and portable applications has accelerated the adoption of distributed systems and cloud computing. To ensure reliability and availability, Site Reliability Engineering (SRE) has become essential. This study compares native cloud monitoring services of Amazon Web Services (AWS), Google Cloud Platform (GCP), and Microsoft Azure within the SRE framework. A PaaS-based sample application was developed using Golang, microservice architecture, API Gateway, and event-driven patterns, supported by relational databases and cloud-native messaging services. Providers’ monitoring tools, along with third-party solutions, were integrated to assess system health. Evaluation focused on objective SRE metrics: variety, cost, alert latency, data collection frequency, and retention. Findings show that GCP provides detailed VM-level metrics without additional configuration in API Gateway setups and achieves the highest message-handling capacity in event-driven architectures. The study offers a comparative analysis of monitoring capabilities and insights into platform suitability for reliability-centric deployments.