Monday, May 14, 2012

An easy way to conduct f-tests on regression coefficients

Suppose you want to conduct a joint test for significance on coefficients of variables that have been expanded in a regression using the xi command. For example suppose you ran the command:

xi: svy, subpop(rual) : y i.lfs i.roofmat  var5 var6 var7 var8

doing an f-test manually on the each of the expanded variables would involve typing (or copying and pasting) the expanded variables, something like:

test _Ilfs_2 _Ilfs_3 _Ilfs_5 _Ilfs_6 _Ilfs_10
test _Iroofmat_2 _Iroofmat_5 _Iroofmat_6 _Iroofmat_7 _Iroofmat_8

As you can see, the numbers on the xi expanded variables do not necessarily increase by one. And this can be cumbersome to type out, or copy and paste, especially if you have many such categorical variables. I'm sure someone has solved this already, but I couldn't find a solution through a web search, so I made my own.

The following program, which you would run after running the reg command, will automatically run f-tests on each group of xi expanded categorical variables. So you would use this program as follows:

xi: svy, subpop(rual) : y i.lfs i.roofmat  var5 var6 var7 var8
easyftest

To use this, just copy the program below and save it as an .ado file in your Stata path to your personal programs directory. The filename should be "easyftest.ado". Let me know if you have any trouble with it. Good luck!

program define easyftest

    local xivars "`_dta[__xi__Vars__To__Drop__]:'"
    local word1 : word 1 of `xivars'
    local pattern = regexr("`word1'","_[0-9]+$","_")
    // di "word1 = `word1'"
    // di "pattern = `pattern'"
    local ftestvars1 "`word1'"
    local count = 0
    local ftestcount 1
    foreach var of local xivars {
        local count = `count' + 1
        if (`count' != 1) {
            local w : word `count' of `xivars'
            // di "w: `w'"
            // check to see whether the next variable is to be included in this list of f-test variables
            if (regexm("`w'","^`pattern'[0-9]+$")) { // there is a match - add this to this list of ftest variables
                // di "pattern match!"
                local ftestvars`ftestcount' "`ftestvars`ftestcount'' `w'"
                // di "ftestvars`ftestcount' : `ftestvars`ftestcount''"
            }   
            else { // no match, create a new list of f-test variables, add this variable to it as the first element, and replace the pattern
                local ftestcount = `ftestcount' + 1
                local ftestvars`ftestcount' "`w'"
                local pattern = regexr("`w'","_[0-9]+$","_")
            }
        }
    }
   
    forv k = 1/`ftestcount' { // Do all the ftest
        // di "ftestvars`k' : `ftestvars`k''"
        // return local ftestvars`k' `ftestvars`k''
        test `ftestvars`k''
    }
    // return scalar N = `ftestcount'

end

Stata tip: Plotting the coefficients estimated from a regression (bar graph in stata)

Suppose you want to make a bar chart/graph/plot of the coefficients (betas) that are returned in the ereturn list from the regression (reg) command. You might want to do this if you want to visualize the relative weight the coefficients give to your estimation. For example, suppose you want to predict consumption based on the assets: car, satellite dish, generator, household size (e.g. if you are working on a Proxy Means Test (PMT) formula). Assume the first three are dummy/binary indicators.

The coefficients estimated from the regression will give you an indication how important each factor is. For example, if the coefficients are, respectively: +5, +1, +3, -15, then you know that the household size dominates the calculation: an additional member reduces predicted consumption more than having all the other assets increases it.

Here is some code for a program (.ado file) that you can call after running the reg command that will create a dataset with the variables in the regression (including the constant) and one observation for each variable, which is the coefficient (see more text after the code below. Yes I know this code is horribly inefficiently written, I just wanted something quick, which means I got something quick and dirty):

program define dataset_coefficients
   
    syntax , Filename(string) [Separator(string)]
    version 9.1
    if ("`separator'" == "") {
        local separator  ","
    }
    // Get the names of the variables to write out. Need to change " o." to " " for making name for the macro to hold the variable labels
    local varnames : coln e(b)
    local coefs ""
    foreach varn of local varnames {
        local coef = _coef[`varn']
        local coefs "`coefs' `coef'"
        local varn1 = regexr("`varn'","o._I","_I")
        if ("`varn'" != "_cons") {
            local varlab_`varn1' : variable label `varn1'
        }
        else {
            local varlab_constant "constant"
        }
    }
    preserve
    drop *
    // Generate the new variable names, and apply the labels
    local variablenamestoplot ""
    foreach varn of local varnames {
        local varn1 = regexr("`varn'","o._I","_I")
        if ("`varn'" != "_cons") {
            gen `varn1' = .
            label var `varn1' `"`varlab_`varn1''"'
            local variablenamestoplot "`variablenamestoplot' `varn1'"
        }
        else {
            gen constant = .
            label var constant "constant"
            local variablenamestoplot "`variablenamestoplot' `constant'"
        }
    }
    // Apply the values to the variables as observations
    set obs 1
    local count = 0
    foreach varn of local varnames {
        local count = `count' + 1
        local coef1 : word `count' of `coefs'
        local varn1 = regexr("`varn'","o._I","_I")
        if ("`varn'" != "_cons") {
            // constant?
            replace `varn1' = `coef1' in 1
        }
        else {
            replace constant = `coef1' in 1
        }
    }
    cap drop __*
    // global variablenamestoplot "`variablenamestoplot'"
    // char [variablenamestoplot] "`variablenamestoplot'"
    notes : `variablenamestoplot'
    save "`filename'" , replace
    restore
end program

After calling this, you can simply load the dataset and graph/chart/plot the coefficients on a bar graph using the following command:

use plotme.dta, clear
 // get the list of variables. I can't just use * because I get some error like __00000 not found. And I don't want to plot the constant.
local listofvars ""
foreach var of varlist * {
        if ("`var'" != "constant") {
            local listofvars "`listofvars' `var'"
        }
}
graph bar (asis) `listofvars', blabel(name, pos(outside) orient(vertical)) legend(off) title("Coefficients ")
graph export coef.png, replace

Let me know how this works for you.

Tuesday, May 8, 2012

Constructing the regression equation with actual coefficients/betas from the e(b) matrix from the ereturn list after running reg with xi and svy

I hope that the title to this post hit all the keywords. So here was my dilema: after running the reg command to estimate regression coefficients (betas), I wanted to apply this equation to a different set of data without having to copy and paste the actual beta hats.

So I have a dataset, hhsurvey.dta, and I estimate the following regression

y = b0 +b1*X1 + b2*X2 + ... bn*Xn

and I get

y_hat
b0_hat
b1_hat
.
.
.
bn_hat

With this, I want to take a different dataset, applicants.dta, with the same variables (but of course different values for these variables), and I want to predict y for the observations in applicants.dta:

y_hat_2 =  b0_hat +b1_hat*X1 + b2_hat*X2 + ... bn_hat*Xn

I could copy and paste the beta_hats from the regression outputs, but this it painful to do even once (I am using many variables because I am using many including categorical variables). Any I suspect I will have to do this many times. My solution was to take the output of the e(b) matrix, which has all the information necessary. After running the regression command:


xi: svy: reg y car i.roofmaterial i.fencematerial i.hhsize ...


you will find some great information stored in the ereturn value "e(b)"


matrix list e(b)


anyways, to make an equation with the regression variables and beta_hats, try the following:

local varnames_rural : coln e(b) // Stores the column names (i.e. variable names) in a local macro.
local equation_rural "" // Will put the equation in this local macro
foreach varn of local varnames_rural { // Loop through all the column (variable) names
    local coef = _coef[`varn'] // This is the beta_hat corresponding to the variable name (inc. categorical vars)
    if ("`varn'" != "_cons") { // The constant in the regression shouldn't be multiplied by anything
        if (`coef' < 0) { // we want to put a "+" before positive coefficients, but not before negative coefficients
            local equation_rural "`equation_rural' `coef'*`varn'"
        }
        else {
            local equation_rural "`equation_rural' + `coef'*`varn'"
        }
    }
    else {
        if (`coef' < 0) {
            local equation_rural "`equation_rural' `coef'"
        }
        else {
            local equation_rural "`equation_rural' + `coef'"
        }
    }
}
di "equation: `equation_rural'"

How about if you want to save this to a file, so that you can load it into a macro in another do file? Try this:

tempname fh
file open `fh' using "myfile.txt", w replace all
file write `fh' "`equation_urban'" _n
file close `fh'

Now, in your new do file that has the applications.dta dataset, with the same variables names, you can use the following code to calcualte y_hat_2 for the applications.dta dataset:

// load the equation
tempname fh2
file open `fh2' using "myfile.txt", r t
file read `fh2' line1
file close `fh2'

di `"line1 = `line1'"'

gen y_hat_2 = `line1'

This should work - leave a comment if it doesn't. Good luck!






Monday, April 30, 2012

Upgrading Perl from 5.12 to 5.14; installing YAML; fixing problem of not being able to install CPAN modules

I don't know what happened since the last post, but I was no longer able to install CPAN modules on my MacBook Pro running Mac OS X Snow Leopard. CPAN didn't like that YAML was not installed, and I wanted to upgrade my Perl installation from 5.12 to 5.14. Finally, I figured all this out. Here is what I did:

1. Install Perl 5.14. Instructions are here. I downloaded the package from http://www.activestate.com/activeperl/downloads, opened the installer and installed the package. This part is easy and straightforward.

After this step, however, my Mac was still using Perl 5.12 rather than the new version. You can check which version your Mac is using by opening a terminal and typing:

perl -v

So you need to add the path to the new Perl version to your path.

2. Instructions for adding the path to the new Perl version to your path your Mac is here. More instructions on how to edit the path file (".profile" file in your user home directory) are available at http://www.tech-recipes.com/rx/2621/os_x_change_path_environment_variable/ and http://www.tech-recipes.com/rx/2618/os_x_easily_edit_hidden_configuration_files_with_textedit/. Doing this tells your Mac where to look for Perl when you type the perl command at the command prompt.

Your .profile file will be located in the "user home directory." For me, this is /Users/shafique. For you it will be whatever /Users/. To check what is in your PATH environment variable, in the terminal window type:

env

or

echo $PATH

Now to add the new path to my path variable, I edited the .profile file using a text editor:

open ~/.profile

Go to the Text Editor, and add the following lines at the end of the text file that the editor just opened:

PATH=/usr/local/ActivePerl-5.14/bin:$PATH
PATH=/usr/local/ActivePerl-5.14/site/bin:$PATH
export PATH
 

Save and close the .profile file. Then go back to your terminal and type:

. ./.profile echo $PATH

Yes, you will type out all three dots on this line. This adds the new Perl path to your environment PATH variable. Now you should be all set. To check which Perl version you are now running, type:

perl -v

and it should show that you are running version 5.14. If not, check your PATH variable:

env

or

echo $PATH

If it doesn't show that the new Perl version directory has been added to your PATH variable, then... well I don't know what to do. 

3. Go ahead an install YAML. In your terminal window, type:

sudo perl -MCPAN -e 'install +YAML'

and then enter your password when the terminal asks for it. 

4. Go ahead and upgrade CPAN. In your terminal window, type:

sudo perl -MCPAN -e 'install CPAN'

5. Uninstall previous versions of Perl:

sudo /usr/local/ActivePerl-5.14/bin/ap-uninstall

6. Remove all symbolic links to previous version of perl in the directories listed in your $PATH variable and replace them with symbolic links to the new versions. Start by finding out which paths may have outdated links:

echo $PATH

Here is what I had to do:

sudo rm /opt/local/bin/perl
sudo rm /opt/local/bin/perl
sudo rm /opt/local/bin/perl5
sudo rm /opt/local/bin/perl5.12
sudo rm /opt/local/bin/perl5.12.4
sudo rm /user/bin/perl
sudo rm /usr/bin/perl
sudo ln -s /usr/local/ActivePerl-5.14/bin/perl /usr/bin/perl
sudo ln -s /usr/local/ActivePerl-5.14/bin/perl /opt/local/bin/perl
sudo ln -s /usr/local/ActivePerl-5.14/bin/perl /opt/local/bin/perl
sudo ln -s /usr/local/ActivePerl-5.14/bin/perl /usr/local/perl

If you don't do this, then your editor (I'm using Komodo Edit) might use the wrong @INC variable, which could cause you problems - e.g. it might think you don't have a cpan module that you actually have installed.

You should be all set now. You should be able to install Perl modules without any problems.

Good luck!

UPDATE (03 Oct 2012): If you're using MacPorts, you can just use the following command:

sudo port install perl5 +perl5_14

But then when you install cpan modules, you have to use macports to install them.







Saturday, April 7, 2012

Trouble installing perl CPAN Modules on Mac OS X

I recently ran into problems trying to install perl CPAN modules. I did the following keyword searches to find answers:

mac won't install CPAN modules

Finally, I decided to upgrade the perl installation on my machine from 5.8 to 5.12. This fixed everything - I can now install perl modules using CPAN without problems. Here is how I did it:

(ref: http://stackoverflow.com/questions/3942520/how-do-i-upgrade-my-macports-perl-installation. And note that I was using macports to do the installation)

sudo port uninstall -f perl5.8
sudo port install perl5 +perl5_12
sudo port -f activate perl5.12  
 
You can check the installation by typing:

perl -v

So far, it works well.

Now, I want to be able to pull context (text) from PDFs. I have come across the following modules:

(ref: http://www.perlmonks.org/?node_id=634794, http://stackoverflow.com/questions/5977969/how-to-parse-pdf-files-in-perl)

PDF::parse
PDF:API2
PDF

I take it that these produce XML from from the PDF content, and then one can parse the XML using the following modules:

(ref: http://stackoverflow.com/questions/5977969/how-to-parse-pdf-files-in-perl)

XML::Twig
XML::Simple

I haven't started pulling PDF content yet though. I think I will just use pdftohtml, like the above link says:

sudo port install pdftohtml

(I could also try xpdf)

Thursday, February 16, 2012

Stata Tutorial 1 (Introduction to using Stata)

...is now up! Click here to view. And please do post comments to let me know what you think.

UPDATE: The font is small in this tutorial - I have made it bigger for tutorial 2. But you can access a google doc with all the commands and output by clicking the following link:

https://docs.google.com/document/pub?id=1IjVOm4-Bvmfm69fgZ7oxvZEYBvcZAgAX2XdiE3WQRbE