Predicting house prices is vital in helping homebuyers, investors, real estate agencies, and policymakers make informed decisions. But it's hard to precisely value a house when there are a number of factors at play – from property details to location to market conditions. The present study accounts for the comparative analysis of four machine learning techniques Multiple Linear Regression (MLR), Support Vector Regression (SVR), Feed Forward Neural Network (FFNN) and Extreme Gradient Boosting (XGBoost)) on predicting house prices by applying the Boston Housing dataset. Root Mean Square Error (RMSE), Mean Absolute Error (MAE), Mean Squared Error (MSE), and Coefficient of Determination (R2) were used to assess the models. Experimental results showed that XGBoost had the lowest RMSE (3.9147), MAE (2.8454), MSE (15.3252) and the highest R2 value (0.7910) among all the models. The second best result was obtained by SVR while FFNN gave an accurate prediction and MLR gave the least accurate result. The analysis results show that the XGBoost model can better capture the complex nonlinear relationship between house attributes and makes it significantly better than the traditional regression model and neural network model. In conclusion, XGBoost is a powerful and stable model that can be used to accurately predict house prices and value real estate.
| Published in |
Science Journal of Business and Management (Volume 14, Issue 3)
This article belongs to the Special Issue Global Challenges in Business: Rethinking Innovations and Sustainability |
| DOI | 10.11648/j.sjbm.20261403.15 |
| Page(s) | 89-97 |
| Creative Commons |
This is an Open Access article, distributed under the terms of the Creative Commons Attribution 4.0 International License (http://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution and reproduction in any medium or format, provided the original work is properly cited. |
| Copyright |
Copyright © The Author(s), 2026. Published by Science Publishing Group |
Machine Learning, Multiple Linear Regression, Feed-forward Neural Network, XGBoost, Support Vector Regression, House Price
Column | Dtype | Description |
|---|---|---|
ZN | float64 | Residential land zone Proportion for lots over 25,000 sq.ft |
RM | float64 | Average rooms per dwelling |
INDUS | float64 | Non-retail business proportion acres per town |
CHAS | float64 | Dummy Charles variable (Tract bounds river 1; 0 otherwise) |
B | float64 | Proportion of blacks by town: 1000(Bk–0.63) ˆ2 |
PTRATIO | float64 | Ratio of pupil–teacher by town |
NOX | float64 | Concentration of Nitric oxides (parts per 10 million) |
AGE | float64 | Prior to 1940 proportion of owner-occupied units built |
RAD | float64 | Accessibility index of radial highways |
CRIM | float64 | Crime rate per capita by town |
TAX | float64 | Property-tax rate per $10,000 |
MEDV | float64 | owner-occupied homes in $1000’: Median value |
LSTAT | float64 | Status of the % lower population |
DIS | float64 | Five Boston employment weighted distances from centers |
CRIM | ZN | MEDV | INDUS | CHAS | NOX | RM | |
|---|---|---|---|---|---|---|---|
Mean | 3.6119 | 11.2119 | 22.5328 | 11.0840 | 0.0700 | 0.5547 | 6.2846 |
Standard Error | 0.3956 | 1.0609 | 0.40886 | 0.3101 | 0.0116 | 0.0052 | 0.0312 |
Median | 0.2537 | 0 | 21.2 | 9.69 | 0 | 0.538 | 6.2085 |
Mode | 0.0150 | 0 | 50 | 18.1 | 0 | 0.538 | 5.713 |
Standard Deviation | 8.720 | 23.389 | 9.1971 | 6.836 | 0.255 | 0.116 | 0.703 |
Sample Variance | 76.042 | 547.040 | 84.5867 | 46.729 | 0.065 | 0.013 | 0.494 |
Kurtosis | 36.568 | 4.133 | 1.49520 | -1.218 | 9.479 | -0.065 | 1.892 |
Skewness | 5.213 | 2.257 | 1.10810 | 0.304 | 3.382 | 0.729 | 0.404 |
Range | 88.969 | 100 | 45 | 27.28 | 1 | 0.486 | 5.219 |
Minimum | 0.0063 | 0 | 5 | 0.46 | 0 | 0.385 | 3.561 |
Maximum | 88.9762 | 100 | 50 | 27.74 | 1 | 0.871 | 8.78 |
Sum | 1755.3708 | 5449 | 11401.6 | 5386.82 | 34 | 280.6757 | 3180.025 |
Count | 486 | 486 | 506 | 486 | 486 | 506 | 506 |
AGE | AGE | DIS | TAX | PTRATIO | B | LSTAT | |
|---|---|---|---|---|---|---|---|
Mean | 68.5185 | 68.5185 | 3.7950 | 408.2372 | 18.4555 | 356.67 | 12.7154 |
Standard Error | 1.2701 | 1.2701 | 0.0936 | 7.4924 | 0.0962 | 4.0586 | 0.3246 |
Median | 76.8 | 76.8 | 3.20745 | 330 | 19.05 | 391.44 | 11.43 |
Mode | 100 | 100 | 3.4952 | 666 | 20.2 | 396.9 | 8.05 |
Standard Deviation | 28.000 | 28.000 | 2.106 | 168.537 | 2.165 | 91.295 | 7.156 |
Sample Variance | 783.973 | 783.973 | 4.434 | 28404.75 | 4.687 | 8334.7 | 51.206 |
Kurtosis | -0.982 | -0.982 | 0.488 | -1.142 | -0.285 | 7.227 | 0.519 |
Skewness | -0.582 | -0.582 | 1.012 | 0.670 | -0.802 | -2.890 | 0.909 |
Range | 97.1 | 97.1 | 10.9969 | 524 | 9.4 | 396.58 | 36.24 |
Minimum | 2.9 | 2.9 | 1.1296 | 187 | 12.6 | 0.32 | 1.73 |
Maximum | 100 | 100 | 12.1265 | 711 | 22 | 396.9 | 37.97 |
Sum | 33300 | 33300 | 1920.291 | 206568 | 9338.5 | 180477. | 6179.7 |
Count | 486 | 486 | 506 | 506 | 506 | 506 | 486 |
Model | Hyperparameter | Value |
|---|---|---|
Multiple Linear Regression (MLR) | Fit Intercept | TRUE |
Copy X | TRUE | |
Positive Constraint | FALSE | |
Support Vector Regression (SVR) | Kernel | RBF |
Regularization Parameter (C) | 100 | |
Gamma | Scale | |
Epsilon (ε) | 0.1 | |
Feed Forward Neural Network (FFNN) | Input Layer Neurons | 5 |
Hidden Layer 1 | 64 neurons | |
Hidden Layer 2 | 32 neurons | |
Hidden Layer 3 | 16 neurons | |
Activation Function | ReLU | |
Output Layer | 1 neuron | |
Optimizer | Adam | |
Learning Rate | 0.001 | |
Batch Size | 16 | |
Epochs | 100 | |
Loss Function | Mean Squared Error | |
XGBoost | Number of Trees (n_estimators) | 100 |
Maximum Tree Depth | 4 | |
Learning Rate | 0.1 | |
Subsample Ratio | 0.8 | |
Column Sampling Ratio (colsample_bytree) | 0.8 | |
Objective Function | reg: squarederror | |
Random State | 42 |
Comparative Models | RMSE | MAE | MSE | R2 |
|---|---|---|---|---|
SVR | 4.984909466 | 3.112876983 | 24.84932239 | 0.661147682 |
FFNN | 5.664001458 | 3.517679092 | 32.08091252 | 0.562535695 |
MLR | 6.33443197 | 3.967722825 | 40.12503 | 0.452843878 |
XGBOOST | 3.914741129 | 2.845376248 | 15.32519811 | 0.791021308 |
MLR | Multiple Linear Regression |
SVR | Support Vector Regression |
FFNN | Feed Forward Neural Network |
XGBoost | Extreme Gradient Boosting |
RMSE | Root Mean Square Error |
MAE | Mean Absolute Error |
MSE | Mean Squared Error |
R2 | Coefficient of Determination |
ICT | Information and Communication Technology |
AI | Artificial Intelligence |
ML | Machine Learning |
LSSVM | Least Squares Support Vector Machine |
PLS | Partial Least Squares |
CFNN | Cascade Forward Neural Networks |
KNN | K-Nearest Neighbors |
| [1] | Adair, A. S., Berry, J. N., & McGreal, W. S. (1996). Hedonic modelling, housing submarkets and residential valuation. Journal of property Research, 13(1), 67-83. |
| [2] | Bin, O. (2004). A prediction comparison of housing sales prices by parametric versus semi-parametric regressions. Journal of Housing Economics, 13(1), 68-84. |
| [3] | Chen, N. (2022). House price prediction model of Zhaoqing city based on correlation analysis and multiple linear regression analysis. Wireless Communications and Mobile Computing, 2022(1), 9590704. |
| [4] |
Singh, P. K., Saraswat, A., Gupta, Y., & Goyal, S. K. (2022). Prediction of Short-Term Solar Radiation Using Machine Learning Methods. In Flexible Electronics for Electric Vehicles: Select Proceedings of FlexEV—2021 (pp. 181-192). Singapore: Springer Nature Singapore.
https://link.springer.com/chapter/10.1007/978-981-19-0588-9_17 |
| [5] | Singh, P. K., Saraswat, A., & Gupta, Y. (2026). Deep learning prediction models for short-term solar photovoltaic power generation forecasting. Next Energy, 11, 100531. |
| [6] | Malang, C. S., Java, E., & Febrita, R. E. (2017). Modeling House Price Prediction using Regression Analysis and Particle Swarm Optimization. International Journal of Advanced Computer Science and Applications, 8(10), 323–326. |
| [7] | Kang, Y., Zhang, F., Peng, W., Gao, S., Rao, J., Duarte, F., & Ratti, C. (2021). Understanding house price appreciation using multi-source big geo-data and machine learning. Land Use Policy, July, 104919. |
| [8] | Greenaway-McGrevy, R., & Sorensen, K. (2021). A Time-Varying Hedonic Approach to quantifying the effects of loss aversion on house prices. Economic Modelling, 99(March), 105491. |
| [9] | Filip F. G., Zamfirescu CB., Ciurea C. (2017) Collaboration and Decision-Making in Context. In: Computer-Supported Collaborative Decision-Making. Automation, Collaboration, & E-Services, vol 4. Springer, Cham. |
| [10] | Aderonke Anthonia Kayode, Noah Oluwatobi Akande, Adekanmi Adeyinka Adegun, Marion Olubunmi Adebiyi (2019), “An automated mammogram classification system using modified support vector machine”, Medical Devices: Evidence and Research, 12, 275-284. |
| [11] | Kayode Anthonia Aderonke, Akande Noah Oluwatobi, Saheed O Jabaru, Oladele O Tinuke (2020), “An Empirical Investigation of the Prevalence of Osteoarthritis in South West Nigeria: A Population-Based Study”, International Journal of Online and Biomedical Engineering (iJOE), 16(1), 100-114. |
| [12] | Helbich, M., Brunauer, W., Vaz, E., & Nijkamp, P. (2014). Spatial heterogeneity in hedonic house price models: The case of Austria. Urban Studies, 51(2), 390-411. |
| [13] | Sean Holly, M. Hashem Pesarana, Takashi Yamagata, (2010). A spatio-temporal model of house prices in the USA", Journal of Econometrics, vol. 158, Issue 1, pp. 160–173. |
| [14] | Mu, J., Wu, F., & Zhang, A. (2014). Housing Value Forecasting Based on Machine Learning Methods. Abstract and Applied Analysis, Volume 2014 (2014), Article ID 648047, 7 pages. Retrieved April 2017, from |
| [15] | Bahia, I. S. (2013). A Data Mining Model by Using ANN for Predicting Real Estate Market: Comparative Study. International Journal of Intelligence Science, 03(04), 162-169. |
| [16] | Sharma, H., Harsora, H., & Ogunleye, B. (2024). An optimal house price prediction algorithm: XGBoost. Analytics, 3(1), 30-45. |
| [17] | Peng, Z., Huang, Q., & Han, Y. (2019, October). Model research on forecast of second-hand house price in Chengdu based on XGboost algorithm. In 2019 ieee 11th international conference on advanced infocomm technology (icait) (pp. 168-172). IEEE. |
| [18] | Zaki, J., Nayyar, A., Dalal, S., & Ali, Z. H. (2022). House price prediction using hedonic pricing model and machine learning techniques. Concurrency and computation: practice and experience, 34(27), e7342. |
| [19] | Zhang, Q. (2021). Housing price prediction based on multiple linear regression. Scientific Programming, 2021(1), 7678931. |
| [20] | Madhuri, C. R., Anuradha, G., & Pujitha, M. V. (2019, March). House price prediction using regression techniques: A comparative study. In 2019 International conference on smart structures and systems (ICSSS) (pp. 1-5). IEEE. |
APA Style
Singh, P. K., Jain, R., Sharma, S. (2026). A Comparative Study of Machine Learning-Based Predictive Models for House Price Prediction. Science Journal of Business and Management, 14(3), 89-97. https://doi.org/10.11648/j.sjbm.20261403.15
ACS Style
Singh, P. K.; Jain, R.; Sharma, S. A Comparative Study of Machine Learning-Based Predictive Models for House Price Prediction. Sci. J. Bus. Manag. 2026, 14(3), 89-97. doi: 10.11648/j.sjbm.20261403.15
AMA Style
Singh PK, Jain R, Sharma S. A Comparative Study of Machine Learning-Based Predictive Models for House Price Prediction. Sci J Bus Manag. 2026;14(3):89-97. doi: 10.11648/j.sjbm.20261403.15
@article{10.11648/j.sjbm.20261403.15,
author = {Praveen Kumar Singh and Rachna Jain and Shikha Sharma},
title = {A Comparative Study of Machine Learning-Based Predictive Models for House Price Prediction},
journal = {Science Journal of Business and Management},
volume = {14},
number = {3},
pages = {89-97},
doi = {10.11648/j.sjbm.20261403.15},
url = {https://doi.org/10.11648/j.sjbm.20261403.15},
eprint = {https://article.sciencepublishinggroup.com/pdf/10.11648.j.sjbm.20261403.15},
abstract = {Predicting house prices is vital in helping homebuyers, investors, real estate agencies, and policymakers make informed decisions. But it's hard to precisely value a house when there are a number of factors at play – from property details to location to market conditions. The present study accounts for the comparative analysis of four machine learning techniques Multiple Linear Regression (MLR), Support Vector Regression (SVR), Feed Forward Neural Network (FFNN) and Extreme Gradient Boosting (XGBoost)) on predicting house prices by applying the Boston Housing dataset. Root Mean Square Error (RMSE), Mean Absolute Error (MAE), Mean Squared Error (MSE), and Coefficient of Determination (R2) were used to assess the models. Experimental results showed that XGBoost had the lowest RMSE (3.9147), MAE (2.8454), MSE (15.3252) and the highest R2 value (0.7910) among all the models. The second best result was obtained by SVR while FFNN gave an accurate prediction and MLR gave the least accurate result. The analysis results show that the XGBoost model can better capture the complex nonlinear relationship between house attributes and makes it significantly better than the traditional regression model and neural network model. In conclusion, XGBoost is a powerful and stable model that can be used to accurately predict house prices and value real estate.},
year = {2026}
}
TY - JOUR T1 - A Comparative Study of Machine Learning-Based Predictive Models for House Price Prediction AU - Praveen Kumar Singh AU - Rachna Jain AU - Shikha Sharma Y1 - 2026/09/08 PY - 2026 N1 - https://doi.org/10.11648/j.sjbm.20261403.15 DO - 10.11648/j.sjbm.20261403.15 T2 - Science Journal of Business and Management JF - Science Journal of Business and Management JO - Science Journal of Business and Management SP - 89 EP - 97 PB - Science Publishing Group SN - 2331-0634 UR - https://doi.org/10.11648/j.sjbm.20261403.15 AB - Predicting house prices is vital in helping homebuyers, investors, real estate agencies, and policymakers make informed decisions. But it's hard to precisely value a house when there are a number of factors at play – from property details to location to market conditions. The present study accounts for the comparative analysis of four machine learning techniques Multiple Linear Regression (MLR), Support Vector Regression (SVR), Feed Forward Neural Network (FFNN) and Extreme Gradient Boosting (XGBoost)) on predicting house prices by applying the Boston Housing dataset. Root Mean Square Error (RMSE), Mean Absolute Error (MAE), Mean Squared Error (MSE), and Coefficient of Determination (R2) were used to assess the models. Experimental results showed that XGBoost had the lowest RMSE (3.9147), MAE (2.8454), MSE (15.3252) and the highest R2 value (0.7910) among all the models. The second best result was obtained by SVR while FFNN gave an accurate prediction and MLR gave the least accurate result. The analysis results show that the XGBoost model can better capture the complex nonlinear relationship between house attributes and makes it significantly better than the traditional regression model and neural network model. In conclusion, XGBoost is a powerful and stable model that can be used to accurately predict house prices and value real estate. VL - 14 IS - 3 ER -