ECON222-Quiz 3 Coverage

Author

SAA

Published

July 1, 2026

Data Used:

For classroom application of this chapter we use the dataset, divisoria_dataset only_v2.xlsx.

Week 5

Previously: Base Line Model - Linear-Linear

In this model, we only used the variables util and sales as regressand and regressor, respectively. Such that:

\[Y_i = \hat{\beta_0}+\hat{\beta_1}X_i+\hat{u_i}\]

Where:

  • \(Y_i\) is utilities expenditure

  • \(X_i\) is level of sales

Regression Results:


Call:
lm(formula = util ~ sales, data = data_1.2)

Residuals:
    Min      1Q  Median      3Q     Max 
-3844.2 -1032.9  -125.6   888.5  5970.2 

Coefficients:
              Estimate Std. Error t value Pr(>|t|)    
(Intercept) -405.95378  831.67365  -0.488    0.627    
sales          0.07420    0.01181   6.281 4.67e-08 ***
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 1844 on 58 degrees of freedom
Multiple R-squared:  0.4048,    Adjusted R-squared:  0.3946 
F-statistic: 39.45 on 1 and 58 DF,  p-value: 4.67e-08

Scatterplot of Lin-Lin

Scatter of Residuals and Sales


Chapter 6: Extension of the Two-Variable Linear Regression Model

Previously, we only considered that the regression model was both linear in the parameter and variable—however, linearity is only strictly applied to the parameter, allowing the variables to be non-linear, we will learn how two-variable models can deal with some interesting practical problems.

6.4 Functional Forms of Regression Models

Remember that we are concerned with models that are linear in parameters, this does not hold for variables.

We will consider commonly usedd models that may be nonlinear in the variables but are linear in the parameters or that can be made so by suitable transformations of the variables.

The following models are as follow:

  • The log-linear model

  • Semilog models

  • Reciprocal models

  • The logarithmic reciprocal model


6.5 The Log-Linear Model (log-log model): How to Measure Elasticity

Consider the following model (exponential regression model)

\[Y_i = \beta_1X_i^{\beta_2}e^{u_i}\]

Recalling the properties of logarithms, and using the natural log \(ln\), we get:

\[lnY_i=ln\beta_1+\beta_2\ lnX_i +u_i\]

Let \(\alpha = ln\beta_1\), then the equation is now:

\[lnY_i=\alpha +\beta_2\ lnX_i + u_i\]

This model is linear in the parameters \(\alpha\) and \(\beta_2\), linear in the logarithms of the variables \(Y\) and \(X\), and can be estimated by OLS regression.

NoteSynonyms for the Model
  • Log-log

  • Double-log

  • Log-linear

If assumptions of the CLRM are fulfilled, the parameters of the equation above can be estimated by the OLS method by letting \(Y_i^* = lnY_i\) and \(X_i^* = lnX_i\).

\[Y_i^* =\alpha +\beta_2 X_i^*+u_i\]

The OLS estimator of \(\hat{\alpha}\) and \(\hat{\beta_2}\) obtained will be the best linear unbiased estimator of (BLUE) \(\alpha\) and \(\beta_2\), respectively.

Feature of the Model

An attractive feature of the log-log model is that the slope of the coefficient \(\beta_2\) measures the elasticity of \(Y\) with respect to \(X\), that is, the percentage change in \(Y\) for a given (small) percentage change in \(X\).

NoteExample

If \(\ Y\) represents the quantity demanded for a good and \(X\) its unit price, \(\beta_2\) measures the price elasticity of demand.

CautionDifference between percent change and a percentage point change

Example:

  • Current unemployment rate is 6%. If this rate were to go to 8%, we say that the percentage point change in unemployment rate is 2%

  • The percentage change in unemployment rate is \(\frac{8\% -6\%}{6\%} = 33\%\)

Special Features

  1. The model assumes that the elasticity coefficient between \(Y\) and \(X\), \(\beta_2\), remains constant throughout, hence the alternative name constant elasticity model.
  2. Although \(\hat{\alpha}\) and \(\hat{\beta_2}\) are unbiased estimators of \(\alpha\) and \(\beta_2\), \(\beta_1\) when estimated as \(\hat{\beta_1} = \text{antilog}(\hat{\alpha})\) is itself a biased estimator.
    • Isn’t that important, we don’t need to worry about obtaining its unbiased estimate.
TipDoes this model fit the data?

In the two-variable model, the simplest way to decide whether the log-log model fits the data is to plot the scattergram of \(lnY_i\) against \(lnX_i\) and see if the scatter points lie approximately on straight line


Class Example

Regression Results


Call:
lm(formula = log(util) ~ log(sales), data = data_1.2)

Residuals:
     Min       1Q   Median       3Q      Max 
-1.15214 -0.20920  0.03578  0.23037  0.81051 

Coefficients:
            Estimate Std. Error t value Pr(>|t|)    
(Intercept)   0.4255     1.7610   0.242     0.81    
log(sales)    0.7141     0.1590   4.493 3.41e-05 ***
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 0.3753 on 58 degrees of freedom
Multiple R-squared:  0.2582,    Adjusted R-squared:  0.2454 
F-statistic: 20.18 on 1 and 58 DF,  p-value: 3.408e-05

Scatterplot of Log-Log

Scatter plot of Residuals and Log of Sales


6.6 Semilog Models: Log-Lin & Lin-Log Models

How to Measure Growth Rate:

Economists, businesspeople, and governments are often interested in finding out the rate of growth of certain economic variables, such as population, GNP, money supply, employment, productivity, and trade deficit.

The Log-Lin Model

Suppose we want to find out the growth rate of personal consumption expenditure on services. Let \(Y_t\) be real expenditure on services at time \(t\) and \(Y_0\) the initial value of the expenditure on services.

Using the compound interest formula, we have:

\[Y_t=Y_0(1+r)^t\]

Taking the natural logarithm \(ln\) of the equation above we get:

\[lnY_t=lnY_0 +t\ ln(1+r)\]

Let \(\beta_1 = lnY_0\) and \(\beta_2 = ln(1+r)\), we have:

\[lnY_t = \beta_1+\beta_2\ t\]

Adding the disturbance term, we obtain;

\[lnY_t=\beta_1+\beta_2\ t+u_i\]

Generally, we can express the equation as:

\[lnY_t=\beta_1+\beta_2\ X_i +u_i\]

  • This model (log-lin) is like any other linear regression model in that the parameters \(\beta_1\) and \(\beta_2\) are linear. The only difference is that the regressand is the \(\text{logarithm of}\ Y\) of the regressor is \(X_i\).

  • This is a semilog model because only one variable is appears in logarithmic form. In this case the regressand, therefore it is called the log-lin model.

Properties of the Log-Lin Model

  • The slope of the coefficient measures the constant proportional or relative change in \(Y\) for a given absolute change in the value of the regressor, that is:

    \[\beta_2=\frac{\text{relative change in regressand}\ {Y_i}}{\text{absolute change in regressor}\ {X_i}} = \frac{\Delta Y \backslash Y}{\Delta X}\]

Multiplying the numerator by 100 will then give the percentage change, or the growth rate, in \(Y\) for an absolute change in \(X\), the regressor. That is 100 times \(\beta_2\) gives the growth rate in \(Y\); 100 times \(\beta_2\) is knows as the semielasticity of \(Y\) with respect to \(X\).

Instantaneous versus Compound Rate of Growth

The coefficient of the trend variable in the growth model, \(\beta_2\), gives the instantaneous (at a point in time) rate of growth and not the compound (over a period of time) rate of growth.

  • The compound rate of growth can be found by taking the antilog of estimated \(\beta_2\) and subtracting 1 from it and multiplying the difference by 100.

  • An illustrative example, the slope of of the coefficient is 0.00705. Therefore:

    \[[\text{antilog}(0.00705)-1]= [e^{0.00705}-1]=0.00707490975\approx0.708\ \text{percent}\]

Linear Trend Model

Researchers sometimes estimate the following model:

\[Y_t=\beta_1+\beta_1\ t + u_i \]

  • Instead of regressing the \(\text{log of}\ Y\) on time, they regress \(Y\) on time, where \(Y\) is the regressand under consideration. This is called a linear trend model, and the time variable \(t\) is known as the trend variable.

    • If the slope of the coefficient is positive, there is an upward trend in \(Y\)

    • If the slope of the coefficient is negative, there is a downward trend in \(Y\)

ImportantChoosing between growth rate model and linear trend model
  • This will depend on upon whether one is interested in in the relative or absolute change—for comparative purposes, it is relative change that is more relevant.

Class Example

Regression Results


Call:
lm(formula = log(util) ~ sales, data = data_1.2)

Residuals:
     Min       1Q   Median       3Q      Max 
-1.14178 -0.19132  0.06373  0.22768  0.67425 

Coefficients:
             Estimate Std. Error t value Pr(>|t|)    
(Intercept) 7.522e+00  1.620e-01  46.438  < 2e-16 ***
sales       1.204e-05  2.301e-06   5.231 2.42e-06 ***
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 0.3592 on 58 degrees of freedom
Multiple R-squared:  0.3205,    Adjusted R-squared:  0.3088 
F-statistic: 27.36 on 1 and 58 DF,  p-value: 2.421e-06

Scatterplot of Log-Lin

Scatterplot of Residuals of Log-Lin


Lin-Log Model

In this model, we want to find the absolute change in \(Y\) for a percentage change in \(X\). Formally:

\[Y_i = \beta_1 +\beta_2\ lnX_i+u_i\]

This is the lin-log model, where \(\beta_2\) is:

\[\beta_2=\frac{\text{Change in}\ Y}{\text{Change in}\ lnX}=\frac{\text{Change in}\ Y}{\text{Relative change in} X}=\frac{\Delta Y}{\Delta X \backslash X}\]

This equation can be rewritten as:

\[\Delta Y=\beta_2(\Delta X\backslash X)\]

This equation states that the absolute change in \(Y\) is equal to the slope times the relative change in \(X\). If the relative change in \(X\) is multiplied by 100, then the equation above gives the absolute change in \(Y\) for a percentage change in \(X\).

Example

If \(\beta_2 = 500\) and \(\frac{\Delta X}{X}=0.01\), then:

\[\begin{aligned} \Delta Y &= \beta_2(\frac{\Delta X}{X})\\ \Delta Y &= 500(0.01)\\ \Delta Y &= 5 \end{aligned}\]

ImportantReminder

Therefore, when \(Y_i=\beta_1+\beta_2\ lnX_i +u_i\) is estimated by OLS, do not forget to multiply the value of the estimated coefficient by 0.01, or, what amounts to the same thing, divide it by 100. If you do not keep this in mind, your interpretation in an application will be highly misleading.

NoteWhen is the Lin-Log Model Useful?

Can be applied in the so-called Engel expenditure models.

  • The total expenditure that is devoted to food tends to increase in arithmetic progression as total expenditure increases in geometric progression.

Class Example

Regression Results


Call:
lm(formula = util ~ log(sales), data = data_1.2)

Residuals:
    Min      1Q  Median      3Q     Max 
-3893.5 -1113.7  -218.4   835.9  6853.4 

Coefficients:
            Estimate Std. Error t value Pr(>|t|)    
(Intercept) -43276.9     9287.2  -4.660 1.90e-05 ***
log(sales)    4323.2      838.3   5.157 3.17e-06 ***
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 1979 on 58 degrees of freedom
Multiple R-squared:  0.3144,    Adjusted R-squared:  0.3026 
F-statistic: 26.59 on 1 and 58 DF,  p-value: 3.173e-06

Scatterplot of Lin-Log

Scatterplot of Residuals and log of Sales


6.8 Choice of Functional Form

Model Equation Slope
\(\left(\dfrac{dY}{dX}\right)\)
Elasticity
\(\left(\dfrac{dY}{dX}\dfrac{X}{Y}\right)\)
Linear \(Y=\beta_1+\beta_2X\) \(\beta_2\) \(\beta_2\left(\dfrac{X}{Y}\right)\)
Log–Linear \(\ln Y=\beta_1+\beta_2\ln X\) \(\beta_2\left(\dfrac{Y}{X}\right)\) \(\beta_2\)
Log–Lin \(\ln Y=\beta_1+\beta_2X\) \(\beta_2Y\) \(\beta_2X\)
Lin–Log \(Y=\beta_1+\beta_2\ln X\) \(\beta_2\left(\dfrac{1}{X}\right)\) \(\beta_2\left(\dfrac{1}{Y}\right)\)

The choice of a particular functional form may be comparatively easy in the two-variable case, because we can plot the variables and get some rough idea about the appropriate model.

The choice becomes harder when we consider the multiple regression model involving more than one regressor. There is no denying that a great deal of skill and experience are required in choosing an appropriate model for empirical estimation. But some guidelines can be offered:

  1. The underlying theory may suggest a particular functional form.
  2. It is good practice to find out the rate of change of the regressand with respect to the regressor as well as to find out the elasticity of the regressand with respect to the regressor. Found in the table above.
  3. The coefficients of the model chosen should satisfy certain a priori expectations.
    • For example, if we are considering the demand for automobiles as a function of price and some other variables, we should expect a negative coefficient for the price variable.
  4. Sometimes more than one model may fit a given set of data reasonably well.
    • In the modified Phillips curve, we fitted both a linear and a reciprocal model to the same data. IN both cases the coefficients were in life with prior expectations and they were all statistically significant. One major difference was that the \(r^2\) value of the linear model was larger than that of the reciprocal model. One may therefore give a slight edge to the linear model over the reciprocal model. But make sure that in comparing two \(r^2\) values the dependent variable, or the regressand, of the two models are the same; the regressor(s) can take any form.
  5. In general one should not overemphasize the \(r^2\) measure in the sense that the higher the \(r^2\) the better the model.
    • \(r^2\) increases as we add more regressors to the model.

    • What is important is the theoretical underpinning of the chosen model,

      • The signs of the estimated coefficients and their statistical significance
    • If the model is good on these criteria, a model with a lower \(r^2\) may be quite acceptable.

  6. In some situations it may not be easy to settle on a particular functional form, in which case we may use the so-called Box-Cox transformations.

Class Example

Then compare them using:

  1. Residual plots (most important for this chapter)

    • Random scatter around zero

    • No curvature

    • No funnel shape (changing variance)

  2. Adjusted \(R^2\) (only compare models with the same dependent variable)

    • Compare Lin–Lin vs Lin–Log (both have \(Y\) as the dependent variable).

    • Compare Log–Lin vs Log–Log (both have \(\ln Y\) as the dependent variable).

    • Do not directly compare the \(r^2\) of a model with $Y$ to one with \(\ln Y\) .

  3. Economic theory

    • Which functional form makes economic sense?

    • Example: Demand and production relationships are often modeled with log transformations because elasticities are meaningful.


Say that we have narrowed down our options for the functional forms to model_linlin and model_loglin. We shall compare their Regression Results


For model_linlin


Call:
lm(formula = util ~ sales, data = data_1.2)

Residuals:
    Min      1Q  Median      3Q     Max 
-3844.2 -1032.9  -125.6   888.5  5970.2 

Coefficients:
              Estimate Std. Error t value Pr(>|t|)    
(Intercept) -405.95378  831.67365  -0.488    0.627    
sales          0.07420    0.01181   6.281 4.67e-08 ***
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 1844 on 58 degrees of freedom
Multiple R-squared:  0.4048,    Adjusted R-squared:  0.3946 
F-statistic: 39.45 on 1 and 58 DF,  p-value: 4.67e-08



For model_loglin


Call:
lm(formula = log(util) ~ sales, data = data_1.2)

Residuals:
     Min       1Q   Median       3Q      Max 
-1.14178 -0.19132  0.06373  0.22768  0.67425 

Coefficients:
             Estimate Std. Error t value Pr(>|t|)    
(Intercept) 7.522e+00  1.620e-01  46.438  < 2e-16 ***
sales       1.204e-05  2.301e-06   5.231 2.42e-06 ***
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 0.3592 on 58 degrees of freedom
Multiple R-squared:  0.3205,    Adjusted R-squared:  0.3088 
F-statistic: 27.36 on 1 and 58 DF,  p-value: 2.421e-06



Plot of the Residuals

Between these two, model_loglin has the most random scatter for residuals.