WEB SERVER CLIENT COMMUNICATION FOR WEB USAGE DATA ANALYSIS USING K-MEANS CLUSTERING | IJET Volume 12 – Issue 5 | IJET-V12I5P29

IJET
International Journal of Engineering and Techniques
ISSN 2395-1303 · Peer-Reviewed · Open Access
📚 Volume 12, Issue 5
📅 September 29, 2026
📄 Pages 266–274
🔖 ID: IJET-V12I5P29

WEB SERVER CLIENT COMMUNICATION FOR WEB USAGE DATA ANALYSIS USING K-MEANS CLUSTERING

Author(s)

V. PRASANNA, PROF. G .SREENIVASULU

Abstract

Information from web server logs can be useful to know about user interaction such as page requests, how often users visit a website, and how long users stay at a website. The actual usage of Web pages, though, can include redundant, incomplete or otherwise irrelevant records that do not lend themselves to direct analysis. This paper introduces a web usage data analysis model which uses the web server-client interaction, data preprocessing, feature extraction and K-Means clustering to find different user behaviour patterns. The suggested workflow involves web usage log collection and preprocessing, followed by the extraction of the behavioral features of web usage (e.g., number of visited pages and web usage time). The experimental analysis was done using a dataset of 15 users, clustered with the K-Means algorithm using 3 clusters. The data was processed, clustered, and visualized using Python, Pandas, NumPy, Matplotlib, and Scikit-learn. These clusters correspond to relatively low, moderate and high activity with the centres of the clusters being 8.4 page visits and 3.8 min, 17.8 page visits and 8.8 min, and 28.4 page visits and 14.0 min, respectively. The outcomes show the feasibility of using K-Means to cluster web users based on their simple usage properties and serve as a starting point for the next step of mining the web usage and analysis of the user behavior.

Keywords

Web Usage Mining, Web Server Logs, Client-Server Communication, K-Means Clustering, User Behaviour, Data Preprocessing, Web Analytics.

Conclusion

In this study, a web usage data analysis framework was proposed, which combined the steps of preprocessing, behavioural features extraction, K-Means clustering and visualization. Web usage mining offers a systematic way to analyze web-server traces and extract useful information from them, and two key steps in the analysis process are pre-processing and the discovery of patterns (Srivastava et al., 2000; Facca and Lanzi, 2005). We analysed 15 users in the experimental implementation with the behavioural features of Page Visits and Time Spent. The results of the K-Means algorithm (3 clusters) were three groups of points with centre coordinates of (8.4 page visits, 3.8 min), (17.8 page visits, 8.8 min) and (28.4 page visits, 14.0 min). The outcomes of these results show that users in the experimental data set can be distinguished by their actual level of use of the website. The principle results of the study are: K-Means can be used to classify web users into clusters based on web usage attributes. The Page Visits and Time Spent metrics are simple representations of the user’s activity and are used for exploratory clustering. The resulting centroids of the clusters give a quantitative description of the user groups identified. The visualization of clusters helps in understanding the difference between the two levels of user activities. Small scale demonstration of the present experiment and the results should not be extended to other web user groups. A more complete web usage analysis could be obtained by using more real-world datasets from server logs in the future and adding more behavioural characteristics. The study proposes the foundation of a computational method for mapping web usage data into meaningful behavioral segments, which is then developed into a practical one.The study thus sets up a basic computational method for mapping web usage data into meaningful behavioral groups which is then developed into a practical one. Additional features in the framework may include more complex log processing, user/session attributes, larger data sets, and comparing and contrasting various methods of log clustering.

References

[1]Cooley, R., Mobasher, B. and Srivastava, J. (1999). Data preparation for mining world wide web browsing patterns. Knowledge and Information Systems, 1(1), 5–32.
[2]Srivastava, J., Cooley, R., Deshpande, M. and Tan, P.-N. (2000). Web usage mining: Discovery and applications of usage patterns from web data. SIGKDD Explorations, 1(2), 12–23.
[3]Facca, F.M. and Lanzi, P.L. (2005). Mining interesting knowledge from weblogs: A survey. Data & Knowledge Engineering, 53(3), 225–241.
[4]Spiliopoulou, M., Mobasher, B., Berendt, B. and Nakagawa, M. (2003). A framework for the evaluation of session reconstruction heuristics in web-usage analysis. INFORMS Journal on Computing, 15(2), 171–190.
[5]Mobasher, B., Cooley, R. and Srivastava, J. (2000). Automatic personalization based on web usage mining. Communications of the ACM, 43(8), 142–151.
[6]Ahmed, M., Seraj, R. and Islam, S.M.S. (2020). The k-means algorithm: A comprehensive survey and performance evaluation. Electronics, 9(8), 1295.
[7]Wu, X., Kumar, V., Quinlan, J.R., Ghosh, J., Yang, Q., Motoda, H., McLachlan, G.J., Ng, A., Liu, B., Yu, P.S., Zhou, Z.-H., Steinbach, M., Hand, D.J. and Steinberg, D. (2008). Top 10 algorithms in data mining. Knowledge and Information Systems, 14(1), 1–37.
[8]Tan, P.-N., Steinbach, M., Karpatne, A. and Kumar, V. (2018). Introduction to Data Mining. 2nd ed. Pearson.

📋 How to Cite This Paper

V. PRASANNA, PROF. G .SREENIVASULU (2026). WEB SERVER CLIENT COMMUNICATION FOR WEB USAGE DATA ANALYSIS USING K-MEANS CLUSTERING. International Journal of Engineering and Techniques, 12(5), 266–274. ISSN: 2395-1303. DOI: https://doi.org/10.5281/zenodo.23032015
© 2026 International Journal of Engineering and Techniques (IJET). All rights reserved. · ijetjournal.org
Submit Your Paper