Introduction
Machine Learning provides powerful techniques for understanding data and making predictions. As part of my internship at Valentius Kryptix, I worked on a practical project focused on Linear Regression and Logistic Regression using Python and Scikit-learn.
The objective of this project was to understand how regression and classification algorithms can be implemented, evaluated, validated, and tuned using real datasets. The project also provided practical experience with data preprocessing, feature scaling, model evaluation, cross-validation, and hyperparameter tuning.
Linear Regression
Linear Regression is a supervised Machine Learning algorithm used to predict a continuous numerical value based on one or more input features.
For this project, I used the Diabetes dataset available through Scikit-learn. The dataset was divided into training and testing sets using an 80:20 split. A Linear Regression model was then trained on the training data and used to make predictions on the test data.
To evaluate the model, I used three important metrics:
- RMSE (Root Mean Squared Error) – measures the average magnitude of prediction errors.
- MAE (Mean Absolute Error) – measures the average absolute difference between actual and predicted values.
- R² Score – indicates how well the model explains the variation in the target variable.
The model achieved an RMSE of 53.85, MAE of 42.79, and an R² score of 0.45 on the test data. I also performed 5-fold cross-validation, which provided a more reliable estimate of model performance.

Logistic Regression
Logistic Regression is a supervised Machine Learning algorithm primarily used for classification problems. Unlike Linear Regression, which predicts continuous values, Logistic Regression predicts the probability of an observation belonging to a particular class.
For this part of the project, I used the Breast Cancer Wisconsin dataset from Scikit-learn to classify observations into malignant and benign categories.
The dataset was divided into training and testing sets using stratified sampling. Since Logistic Regression can be affected by differences in feature scales, I applied StandardScaler to standardize the input features before training the model.
The baseline Logistic Regression model achieved approximately 98.25% accuracy, with precision, recall, and F1-score also around 98.61% on the test set.

Model Evaluation and Tuning
Model evaluation was an important part of the project. In addition to accuracy, I analysed precision, recall, F1-score, and the confusion matrix to better understand classification performance.
I also used GridSearchCV to perform hyperparameter tuning for Logistic Regression. The best cross-validation parameter was C = 0.1, with a cross-validation accuracy of approximately 98.02%.
An important learning from this process was that hyperparameter tuning does not always improve performance on a particular unseen test set. In my experiment, the tuned model achieved slightly lower test accuracy than the baseline model. This highlighted the importance of comparing cross-validation results with independent test-set performance rather than assuming that tuning will always produce better results.

Key Learnings
This project helped me develop a stronger practical understanding of Machine Learning workflows.
Some of my key takeaways were:
- Understanding the difference between regression and classification problems and selecting an appropriate algorithm.
- Learning the importance of feature scaling, cross-validation, and suitable evaluation metrics.
- Understanding how hyperparameter tuning works and how model performance should be evaluated on unseen data.
- Learning how to interpret model coefficients to understand the influence of different features.
Conclusion
Working on Linear and Logistic Regression provided valuable hands-on experience in implementing Machine Learning models using Python and Scikit-learn. The project covered the complete workflow from dataset preparation and model training to evaluation, validation, visualization, and hyperparameter tuning.
This experience has strengthened my understanding of practical Machine Learning and given me a better foundation for working with more advanced algorithms and real-world data science projects.
Project:
GitHub: https://github.com/Akash-Yadav90/DS-5_Linear_Logistic_Regression_Akash_Kumar.ipynb



Leave a Reply
You must be logged in to post a comment.