Supervised Learning Algorithms: Linear and Logistic Regression
In supervised learning, models learn from labeled training data to map inputs to desired outputs. Two of the most fundamental algorithms for solving these problems are linear regression and logistic regression.
While both use linear combinations of input features, they serve completely different purposes: linear regression predicts continuous numerical values, whereas logistic regression handles binary classification tasks.
Linear Regression for Continuous Outputs
Linear regression models the relationship between a dependent continuous target variable ($y$) and one or more independent features ($x$) by fitting a straight line, known as the line of best fit.
from sklearn.linear_model import LinearRegression
# Initialize and train model
lin_reg = LinearRegression()
lin_reg.fit(X_train, y_train)
# Predict continuous values
predictions = lin_reg.predict(X_test)
The algorithm minimizes the Residual Sum of Squares (RSS) between the actual data points and the predicted regression line using Ordinary Least Squares (OLS).
Logistic Regression for Binary Classification
Despite its name, logistic regression is a classification algorithm rather than a regression model. It predicts the probability that a given input belongs to a specific category (e.g., 0 or 1, true or false, spam or not spam).
from sklearn.linear_model import LogisticRegression
# Initialize and train model
log_reg = LogisticRegression()
log_reg.fit(X_train, y_train)
# Predict class probabilities and labels
probabilities = log_reg.predict_proba(X_test)
predictions = log_reg.predict(X_test)
Logistic regression passes the linear combination of inputs through a sigmoid (logistic) function, compressing the output values to a range between 0 and 1.
Linear vs. Logistic Regression Comparison
Understanding when to apply each algorithm is crucial for effective machine learning model design:
• Output Nature: Linear regression outputs a continuous real number, whereas logistic regression outputs a discrete probability bounded between 0 and 1.
• Activation Function: Linear regression uses a direct linear equation ($y = mx + c$), while logistic regression applies the sigmoid activation function to map linear outputs to probabilities.
• Use Cases: Linear regression is ideal for forecasting trends like house prices or temperature, while logistic regression excels at classification problems like medical diagnosis or churn prediction.
Summary
Linear and logistic regression form the backbone of supervised machine learning. Mastering these two algorithms provides the fundamental intuition required to understand more complex models like neural networks and gradient boosting.