Are countries that are more economically diverse more developed? In modern economics, the answer is generally yes — countries that are overly reliant on one product or industry often face economic downturn as all other industries struggle. This project uses data from the Harvard Atlas of Economic Complexity to empirically determine the veracity of that common theory. When counting the raw number of products a country exports competitively, or adjusting for the complexity (which adjusts for how "niche" products are on the global stage), there is a strong positive correlation with GDP. However, when counting the number of competitively-exported products reweighted by similarity between products themselves, there is a strong negative correlation with GDP, suggesting that though richer countries tend to produce more products and more complicated ones, the actual diversity of their exports decreases as they develop. After determining that the complexity of countries' products Granger-causes GDP using the Emirmahmutoglu-Kose procedure, I concluded that though it is important for countries to expand into more products ("diversify") if they want to grow, that expansion cannot be truly diverse but must be directed specifically into nearby, more complex products.
Introduction
Introduction
It is a common economic theory that countries which export more diverse products tend to be more developed. Though early economists such as David Ricardo supported the idea that nations would become more prosperous by specializing in certain products (those which they had a comparative advantage in), more recent theories and empirical evidence have supported a different interpretation. For example, modern Economists point as an example to the so-called "Resource Curse" — the idea that countries that have been "blessed" by an abundance of natural resources or oil deposits, the extraction of which naturally becomes the overwhelming majority of that country's economy, are generally poorer, less developed, and more prone to corruption. Similarly common in modern discussion is "Dutch Disease," the process in which a country's over-reliance on one highly lucrative sector appreciates the country's currency, making it harder for all of the country's other industries' producers to export and thus damaging the economy.
The goal of this project is to either confirm or refute that claim — that diversity goes with development — using real-world data. In particular, I use historical trade data classified by product type for each country to determine if countries whose export sector is more fragmented are generally richer or poorer. Section 2 ("Data Import") discusses the data sources and raw metrics used in this study. In Section 3 ("Calculating Metrics for Diversity and Complexity"), I lay out three possible metrics for "diversity" based on three separate definitions of the word. Section 4 ("Econometric Models") fits these calculated metrics to values of GDP per capita for each country using Linear OLS regressions to determine a model that best explains economic growth. Finally, Section 5 ("Determining Causal Inference and Granger Causality") uses the results from the previous sections, along with more econometric tests, to determine whether a causal relationship exists between diversity and GDP, concluding with implications for policymakers and the design of targeted economic development strategies.
The goal of this project is to either confirm or refute that claim — that diversity goes with development — using real-world data. In particular, I use historical trade data classified by product type for each country to determine if countries whose export sector is more fragmented are generally richer or poorer. Section 2 ("Data Import") discusses the data sources and raw metrics used in this study. In Section 3 ("Calculating Metrics for Diversity and Complexity"), I lay out three possible metrics for "diversity" based on three separate definitions of the word. Section 4 ("Econometric Models") fits these calculated metrics to values of GDP per capita for each country using Linear OLS regressions to determine a model that best explains economic growth. Finally, Section 5 ("Determining Causal Inference and Granger Causality") uses the results from the previous sections, along with more econometric tests, to determine whether a causal relationship exists between diversity and GDP, concluding with implications for policymakers and the design of targeted economic development strategies.
Data Import
Data Import
Time Series data of Country Exports (nominal USD) by Product Type
Time Series data of Country Exports (nominal USD) by Product Type
The Harvard Atlas of Economic Complexity provides extensive data regarding country exports. The specific dataset I imported was divided by year, country of export, and product type. The product classification was provided to very high granularity, and I used the HS 1992 six-digit codes for my analysis.
A note on the HS Codes
A note on the HS Codes
Import the Goods Data (time series, by country, by HS1992 Good ID)
In[]:=
goodsImport=CloudGet[CloudObject[
https://www.wolframcloud.com/obj/28jacobc/goodsImport
]];baseCountryGoodsData=RenameColumns[ConstructColumns[goodsImport,{"country_iso3_code","product_id","product_hs92_code","year","export_value"}],{"export_value"->"exports","country_iso3_code"->"country","product_hs92_code"->"productString","product_id"->"product"}]Out[]=
Tabular
Extract a list of all countries and their populations from the goods data:
In[]:=
countryList=Normal@DeleteDuplicates[baseCountryGoodsData[All,"country"]];populationList=If[NumberQ[Quiet[QuantityMagnitude[CountryData[#,"Population"]]]],QuantityMagnitude[CountryData[#,"Population"]],0]&/@countryList;populationAssoc=Association@@Rule@@@Transpose[{countryList,populationList}];countryList[[1;;10]]
Out[]=
{AFG,ALB,ATA,DZA,ASM,AND,AGO,ATG,AZE,ARG}
To clean the data, I started by removing countries in the dataset with a population less than 100,000. For these regions, the low population (and thus significantly smaller economy) caused the export data to be highly volatile. Even just one large shipment of a product by a single corporation within that country could greatly increase that product's prevalence for that year. Additionally, many of these places were already not countries to begin with, but rather small, offshoot regions of larger nations (e.g. Åland Islands, Bouvet Island, Gibraltar).
Find listed countries with a population less than or equal 100,000:
In[]:=
banned=Join[Keys[Select[populationAssoc,#<=100000&]],{"ANS","USP"}]
Out[]=
{ATA,ASM,AND,ATG,BMU,BVT,VGB,CYM,CXR,CCK,COK,DMA,FRO,FLK,SGS,ATF,PSE,GIB,KIR,GRL,GRD,HMD,VAT,MSR,NRU,ANT,ABW,SXM,BES,NIU,NFK,MNP,FSM,MHL,PLW,PCN,BLM,SHN,KNA,AIA,SPM,VCT,SMR,SYC,TKL,TON,TCA,TUV,ANS,ANS,USP}
Update countryList to remove the banned countries:
In[]:=
countryList=DeleteElements[countryList,banned];
Create one base tabular for all product data, removing countries with population less than 100,000 people and rows representing products with 0 exports:
In[]:=
baseCountryProductData=Discard[Sort[Select[Discard[baseCountryGoodsData,#exports==0&],#year>=1995&]],MemberQ[banned,#country]&];
Get list of products:
In[]:=
productList=Normal@DeleteDuplicates[baseCountryProductData[All,"product"]];countryAssoc=AssociationThread[countryList,Range[Length[countryList]]];productAssoc=AssociationThread[productList,Range[Length[productList]]];
Creates an association to map HS 1992 Codes (used specifically for goods) to the product ID (used by the Harvard Growth Lab):
In[]:=
hsCodes=Association@@(#[[2]]->#[[1]]&/@Normal@DeleteDuplicates[ConstructColumns[baseCountryProductData,{"productString","product"}]]);
Time Series data of Country GDP (nominal USD) and Population
Time Series data of Country GDP (nominal USD) and Population
All other data was imported from the World Bank Open Data Catalogue, which provides macroeconomic datasets regarding common metrics. In my case, I imported data for country GDP and country Population.
Imports the Time Series nominal GDP data for each country:
In[]:=
Data Cleaning
Data Cleaning
Import Time series population data:
In[]:=
Data Cleaning
Data Cleaning
Extracts an association of each country to its GDP based on a given year:
In[]:=
gdpYear[year_]:=Apply[Association,#1->If[MissingQ[#2],0,#2]&@@@Normal@ConstructColumns[gdpData,{"Country Code",N[year]}]];gdpYear[2024][[1;;10]]
Out[]=
ABW4.16759×,AFE1.25256×,AFG1.77785×,AFW7.37367×,AGO1.03081×,ALB2.70375×,AND4.04425×,ARB3.74504×,ARE5.52325×,ARG6.38365×
9
10
12
10
10
10
11
10
11
10
10
10
9
10
12
10
11
10
11
10
Does the same for the population:
In[]:=
popYear[year_]:=Apply[Association,#1->If[MissingQ[#2],0,#2]&@@@Normal@ConstructColumns[popData,{"Country Code",ToString[year]<>".0"}]];popYear[2024][[1;;10]]
Out[]=
ABW107995.,AFE7.69281×,AFG4.26475×,AFW5.20655×,AGO3.78858×,ALB2.37713×,AND81938.,ARB4.92613×,ARE1.09864×,ARG4.56962×
8
10
7
10
8
10
7
10
6
10
8
10
7
10
7
10
Using the ECI and gdpYear associations, creates an association of the GDP per capita of each country for any given year:
In[]:=
gdpPerCapitaRaw[year_]:=Association@@Rule@@@Transpose[{countryList,Divide@@@Transpose[{gdpYear[year][#]&/@countryList,popYear[year][#]&/@countryList}]}]gdpPerCapitaRaw[2024][[1;;10]]
Out[]=
AFG416.871,ALB11374.,DZA5752.99,AGO2720.82,AZE7294.64,ARG13969.8,AUS64610.,AUT58268.9,BHS39455.4,BHR29717.1
Time Series Data of Export to GDP ratio by Country
Time Series Data of Export to GDP ratio by Country
Import the Export to GDP ratio time series data:
In[]:=
Out[]=
Tabular
In[]:=
etGDP2024=Association@@Rule@@@Transpose[{Normal@etGDPdata[[All,2]],Normal@etGDPdata[[All,32]]}]/._Missing->Indeterminate;
Calculating Metrics for Diversity and Complexity
Calculating Metrics for Diversity and Complexity
"Diversity" is inherently a difficult word to quantify. In this project I used three different measures of "diversity," each based on a different interpretation of what the diversity metric should reflect.
In[]:=
A Basic Diversity Index
A Basic Diversity Index
Perhaps the simplest way to measure diversity is just by asking, "how many products does this country produce?" That is, if Country i produces n products competitively on the global scale, then the diversity =n.
Δ
i
In[]:=
Length[productList]
Out[]=
5039
There are roughly 5,000 unique products in the dataset (which, in turn, was defined from the HS 1992 six-digit classification). This means that by this definition of diversity, the maximum possible number of products a country could produce is capped around 5,000.
Function to create a matrix of the raw dollar value of exports by Country and Product
In[]:=
createPCMatrix[year_]:=Module[{filtered,finalMatrix,sumRows,sumColumns,sumWorld},filtered=Normal[Select[baseCountryProductData,#year==year&]];finalMatrix=ConstantArray[0,{Length[countryList],Length[productList]}];(finalMatrix[[countryAssoc[#["country"]],productAssoc[#["product"]]]]=#["exports"])&/@filtered;finalMatrix]
Defining how many products a country exports competitively is a bit trickier. One possible solution is to use the Revealed Comparative Advantage (RCA); the process for calculating it is shown below:
1
.define a matrix with rows being defined by countries and columns being products. Each entry reflects the total dollar value of a country's exports in a certain year for that specific product.
M
cp
2
.calculate the percent of the country 's exports that are comprised by that specific product.
c
3
.calculate the percent of the world's exports that are comprised by that specific product.
4
. divide the result from step #2 by step #3. Symbolically,=÷∑÷∑.
RCA
cp
M
cp
M
c
M
wp
M
w
For example, a country with an RCA of 0.05 is exporting one twentieth the amount it should to match the global average for each country (adjusted for export size). A country with RCA < 1 is not pulling its weight on a product; a country with RCA > 1 is exporting more than it would if all countries exported a uniform basket corresponding to the global distribution.
Function to calculate the Revealed Comparative Advantage (RCA) of any given country and product (zero if N/A)
In[]:=
calculateRCA[countryProduct_,country_,product_,world_]:=If[countryProduct==0||country==0||product==0||world==0,0,(countryProduct÷country)/(product÷world)//N]
For later steps, my RCA matrix was with a modified value based on the originally calculated RCA values. I used , so 0.5 reflected a country exactly pulling its weight.
RCA
1+RCA
Function to create a C/P matrix with modified RCA values
Once I had a matrix with all possible RCA values for each country-product pair, I binarized it (rounding up to 1 if the country has a comparative advantage in a certain product and down to 0 if it doesn't).
Sums the rows of a binarized RCA Matrix showing how many products each country has a comparative advantage in:
To calculate the diversity value, then, I just counted the number of unique products a country has a comparative advantage in (taking the total of one row of the binary matrix).
Plot a histogram of all countries' basic diversity values for the year 2024:
Plotting those values on a histogram (above), there was a significant right skew in the data, suggesting that an elite few countries export many thousands of unique products. To make this data more symmetrical, I took the common log of the diversity count and standardized the ensuing distribution to calculate the basic diversity metric.
Creates function to extract the countries which were not removed by the data cleaning (for each year)
Creates a list/association of Log[diversity]
The resulting distribution was much more normal:
I tested this data with two countries of vastly different export structures. Venezuela (VEN) is highly specialized, with more than 80% of the entire country's exports being in HS product 2709.00 (crude oil). By contrast, Sweden (SWE) produces multiple different types of products across sectors.
In this standardized model, Sweden has a highly diverse (Δ = 1.19) economy, whereas Venezuela does not (Δ = -1.14).
Using Rao-Stirling Quadratic Entropy
Using Rao-Stirling Quadratic Entropy
Redefines the association of countries to remove countries which were too small or insignificant:
Redefines the association of products:
Get a vector of all product exports by one country in a certain year:
The first step was to find a way to quantify how "similar" two products were. To do this, I found the correlation matrix (the "product space") of the RCA Matrix found in the previous section, since similar products would naturally be produced more often together.
Defines a correlation matrix (the product space) for the RCA matrix:
Plots a matrix of the product space in 2024:
What I was looking for was not how similar two products were, but actually how different they were. I inverted the product space matrix to get a matrix of the dissimilarity between products:
In order to assign a numerical value of "diversity" to a country's different RCA values within the product space, I used a version of the Rao-Stirling index, originally used for measuring biodiversity within an ecosystem while adjusting for how similar different species were. In my case, the base idea was the same (only I used it for economic diversity rather than biodiversity):
Define the Rao-Stirling quadratic entropy index for Econometrics:
Plot a histogram of standardized Rao values in 2024:
Economic Complexity Index
Economic Complexity Index
The third (and last) main metric of diversity that I looked at was the Economic Complexity Index (ECI), first laid out in the seminal paper by Hidalgo and Hausmann in 2009. The ECI looks not only at the products a country produces, and how different they are, but it also weights by how "ubiquitous" (niche) the products that a country produces are. If everyone else in the world also produces soybeans, for example, a nation's soybeans export will not be worth as much, since soybeans are not a "complex" product. In this way, the ECI combines both diversity (that is, how varied a country's exports are) but also ubiquity (how important those exports are on a global scale). Since ECI is defined off of PCI and PCI is defined off of ECI, both indices are created through a recursive method of reflections. A simple example of the ECI calculation is shown below.
Defines functions to query diversity (how many products does a country make) and ubiquity (how many countries make a product)
Defines function to modify ECI based on PCI, then modify PCI based on new ECI
Creates function to find the ECI and PCI values by country, outputting in the form {ECI, PCI}
Plots a histogram of the ECI values for all countries in 2024:
View a sample association of ECI values:
Econometric Models
Econometric Models
I now had multiple variables which I wanted to test the relationships between. Those variables (and what I abbreviated them as) are listed below:
◼
Basic Diversity Index (BDI) — Measures the raw number of products that a country has a comparative advantage in, then takes the common log of that value.
◼
Economic Complexity Index (ECI) — Measures the weighted mean of a country's product complexities (measured through the product complexity index, PCI), weighted by the product RCA values for that country. First proposed by Hidalgo and Hausmann in 2009.
◼
Rao-Stirling Quadratic Entropy (RAO) — Measures the diversity of a country's RCA product exports, penalizing for heightened similarity between products
Linear Regression with the Basic Diversity Index & GDP per capita
Linear Regression with the Basic Diversity Index & GDP per capita
My goal for this section was to determine whether the raw number of products that a country exported (one possible measure of diversity) correlated with GDP per Capita.
Redefine GDP per capita to only include active countries:
Get GDP per capita to Basic Diversity data as a list, and plot its values in 2023:
Since GDP per capita seemed to have a heavy right tail, I took logarithms of its values:
Redefine GDP per capita to Diversity as a Log-Log model, and plot histogram of Log[GDP per capita] in 2023:
Plot a new ListPlot of BDI to log(GDP):
Visually, there appeared to be a positive correlation. To confirm this, I took a linear regression of the data:
Express the linear regression for BDI and GDPcap in Mathematica:
Plot the fit with the BDI model datapoints:
The linear model fit returned a significant (p << 0.001) value for the slope of 0.602, suggesting that for every 10% increase in the number of products competitively produced by a country, the GDP per capita increases by 6.02%. To make sure the model was valid, I next tested for some of the assumptions. In particular, I wanted to check for homoscedasticity and the normality of residuals:
Plot residuals to test for homoscedasticity:
Because the dots appeared to be randomly distributed above and below the line, I concluded that this model did exhibit homoscedasticity.
Plot Q-Q plot to test for the normality of residuals:
Because the Q-Q plot largely followed the line y=x throughout most of the plot, I concluded that the residuals of this basic model did exhibit normality (though they had slightly heavy tails).
Linear Regression with ECI & GDP per capita
Linear Regression with ECI & GDP per capita
The next metric whose model I considered was ECI, since it not only considered diversity but also the ubiquity of each product a country produced.
Get GDP per capita (log.) to ECI for model data:
Apply a linear model fit for GDP per capita and ECI
Plot the ECI model fit in 2023:
Plot the Residuals of the ECI 2023 model:
The GDP per capita to ECI also had a significant (p << 0.001) slope of 0.464, which suggests that each 0.1 value increase of ECI increases a country's GDP per capita by roughly 4.64%. However, the residuals on the residual plot did appear slightly U-shaped, so the next model I tried was a quadratic fit.
Linear Regression with (ECI, ECI^2) & GDP per capita
Linear Regression with (ECI, ECI^2) & GDP per capita
Define the quadratic model to try to correct for the U-shape in the residual plot:
Plot the quadratic model and its datapoints in a scatterplot:
Visually, the quadratic fit does seem to be more effective, but to make sure, I ran a partial F-test:
Define function for F-test and run using two ECI models (quadratic and linear)
Because the p-value was significant ( < 0.001), I rejected the null hypothesis and concluded that the full (quadratic) model had better explanatory power for the data than the reduced (linear) model.
Linear Regression with Rao diversity index & GDP per capita
Linear Regression with Rao diversity index & GDP per capita
The last metric I considered was the Rao diversity index, which isolated the diversity while still accounting for similarities between products. At this point, I had already run the other two metrics and was fully expecting the third to also show a strong positive correlation.
Get GDP per capita (log.) to Rao Diversity for model data:
Create model:
Instead, what the model and data showed was a significant (p << 0.001) negative correlation between diversity (adjusted for product distance) and GDP per capita. Though surprising, and originally counterintuitive when comparing the fit to that of the ECI data, this data can still be explained by a country's position in the product space. Countries that are more developed, such as the U.S. or China, produce many different products across different sectors, thus giving the appearance of "diversity" of exports. However, across all of these sectors, the products which these countries export are complex and technological: cars, computer chips, advanced electronics, etc.; as such, even though they export more products, those products are all actually quite similar in their means of production (that is, each product's production involves complex machinery). By contrast, poorer countries generally export a wide variety of different natural resources, such as oil, cereal agriculture, or stone/minerals. These products are generally not as advanced, but the means of production for creating them are still very different (and thus, by the Rao index, those countries are also very diverse).
This paradigm can be viewed in a graphical representation of the product space:
This paradigm can be viewed in a graphical representation of the product space:
Classify full six-digit HS 1992 codes into their 2-digit chapters:
Define descriptions for HS 1992 2-digit chapters:
Define colors for the HS 1992 2-digit chapters:
Define colors/HS 1992 numbers for each row of the product space matrix:
Use dimension reduction techniques to visualize the product space:
In this depiction of the product space, the two axes both convey much information. The x-axis appears to show product complexity in general, with complex, technologically advanced, and less diverse products existing on the left side of the space and more basic and less diverse natural resource-oriented products existing on the right side of the space. The y-axis appears to show degree of capital or resource intensity, with capital-intensive products existing at the top of the space and labor-intensive products existing at the bottom.
Richer, more developed countries tend to occupy the space to the left of the graph — that space tends to be filled with many complex machineries and electronics, which are not only advanced but similar in production to each other. In contrast, poorer countries tend to occupy the space to the right of the graph, which is filled with disparate natural resources with little relationship to each other.
Richer, more developed countries tend to occupy the space to the left of the graph — that space tends to be filled with many complex machineries and electronics, which are not only advanced but similar in production to each other. In contrast, poorer countries tend to occupy the space to the right of the graph, which is filled with disparate natural resources with little relationship to each other.
Plot a given country's presence in the product space:
Multiple Linear Regression with all independent variables and an interaction term
Multiple Linear Regression with all independent variables and an interaction term
Define all the data for the full model:
Define the full model:
Test the Variance Inflation Factors on the full model:
Extract the parameter estimates for a reduced model:
Calculate VIF values for the full model:
Because all other parameters' p-values were now less than 0.05, I ended the backwards selection process.
The coefficient of the interaction term between ECI and the Exports-To-GDP ratio was very small (close to zero), suggesting that there is no change in the predictive power of ECI depending on countries' reliance on exports.
Determining Causal Inference and Granger Causality
Determining Causal Inference and Granger Causality
In econometrics, determining true causality is often extremely difficult or impossible; the next best choice is "Granger causation," which asks if one time series variable has predictive power in estimating the future value of another time series variable. Therefore, my goal for this section was to try to determine if there existed any causality between ECI (which had the strongest fit, among the three metrics in the previous section) and GDP per capita. I started by just finding/computing the time series data for both ECI and GDP per capita, and plotting.
Create function to query time series ECI data by country:
For most countries, their ECI value stayed largely constant, since the index is naturally standardized:
Plot ECI values over time for USA, Great Britain, and Germany:
For others, their ECI value fell or rose dramatically as the country's economic system and production changed over the years:
Plot ECI values over time for Vietnam, Venezuela, and Cambodia:
Create function to query time series GDP per capita data by country:
Plot the USA GDP per capita as a function of time:
One requirement to directly test for Granger Causality is for the data to be stationary. I tested for that using the Augmented Dickey-Fuller test (ADF):
Calculate the ADF p-value of the GDP per capita time series for a set of test countries:
Calculate the ADF p-value of the ECI time series for a set of test countries:
Across both the ECI and GDPcap data, the data was not stationary (p > 0.05). One solution to correct for this is first order differencing (taking the derivative of the time series graphs and seeing if the derivatives themselves are stationary).
Calculate the ADF p-value of the first-order-differenced GDP per capita time series for a set of test countries:
Calculate the ADF p-value of the first-order-differenced ECI time series for a set of test countries:
Apart from the seemingly sole exception of the United States (which had a p-value quite close to 0.05 anyway), nearly every country exhibited stationarity in its first-order differenced data, both for ECI and GDP per capita. This suggested that both time series were first-order integrated. The next logical step, then, was to ask if they were also cointegrated — that is, if they shared a long-term, mean-reverting relationship with each other. To do this, I tested using the Engle-Granger Test by running a standard OLS regression between the ECI and GDP per capita data with datapoints in each year, and determining if the OLS model's residuals were stationary using the ADF test.
Run the Engle-Granger Test for cointegration for a set of countries:
For every country, the p-value for the Engle-Granger test was significantly below the significance threshold of 0.05, so I concluded that the time series data for ECI and GDP per capita were cointegrated with order 1. From this result I then concluded:
◼
Running the basic Granger Causality test on the first-order difference of these variables would not work, since it would induce misspecification and omit the long-run equilibrium relationship inherent in the system.
And, arguably more importantly,
◼
By the Granger Representation Theorem, because these two time series sets of data were cointegrated, there must exist causality in at least one direction. That is, it must be so that either ECI granger causes GDP, or GDP granger causes ECI, or both.
In order to determine the direction of causality as well, I used a combination of two econometric methods: the Toda Yamamoto procedure corrects for cointegration in small time series data by deliberately overfitting the data, and the Dumitrescu-Hurlin procedure helps combine the Wald Statistics of panel causality data into one singular metric for the entire dataset. The two procedures were combined in a 2011 paper by Emirmahmutoglu and Kose.
Define a helper function to get the last p years of data for any given time:
Testing Granger Causality of ECI on GDP
Testing Granger Causality of ECI on GDP
Create the GDP autoregression function:
Test the GDP autoregression function with a country and lag value:
In order to determine the optimal lag length, I used the AICc critereon across a range of different countries.
Plots the AICc values for the ECI GDP model for a variety of countries:
In general, the minimum value was achieved around k=4, so that was the number of lags that I used.
In order to run the Todo-Yamamoto part of the E-K test, I needed to calculate the Wald statistic for the multivariate autoregression:
Create function to calculate the Wald statistic for a given country:
View Wald statistics for certain countries:
Get list of Wald Statistics for all countries:
Define N (number of samples):
Define T (30 time periods – 4 lags):
Calculate the Modified Z-statistic:
In the E-K test, this modified Z-statistic is compared to the Normal Distribution:
With a significant p-value (p << 0.001), I rejected the null hypothesis and concluded that ECI does Granger-cause GDP per capita in at least one country.
Testing Granger Causality of GDP on ECI
Testing Granger Causality of GDP on ECI
Create the ECI autoregression function:
In order to determine the optimal lag length, I again used the AICc criterion across a range of different countries:
Plots the AICc values for the GDP ECI model for a variety of countries:
In general, the minimum value was achieved at k=1, so that was the number of lags that I used.
Create function to calculate the wald statistic for a given country:
Compare the Wald Statistic to the Chi Square distribution with 1 degree of freedom:
View Wald Statsitics for a set of countries:
Get list of Wald Statistics for all countries:
Define N:
Define T:
In the E-K test, this modified Z-statistic is compared to the Normal Distribution:
With an insignificant p-value (p > 0.05), I accepted the null hypothesis and concluded that GDP does not Granger-cause ECI in any country.
Since past-values of ECI can effectively predict future values of GDP, while past values of GDP cannot do the same for ECI, there is one-way granger causality. This suggests that countries which are richer will not necessarily become more complex, but those who are complex (while still having a low GDP) will see their GDP per capita values rise rapidly to catch up to the high ECI.
Conclusion
Conclusion
This study evaluated the empirical association between export diversification and macroeconomic development across a panel of 173 countries from 1995 to 2024. While classical trade theory champions specialization, modern data reveals a more nuanced dynamic. By implementing three distinct structural metrics — a Basic Diversity Index, Rao-Stirling Quadratic Entropy, and the Economic Complexity Index (ECI) — I gained a more nuanced picture of how diversity goes with growth. Though the Basic Diversity Index and Economic Complexity Index both correlated positively with GDP per capita in each year, The Rao-Stirling index correlated negatively; this suggests that countries which are more developed generally produce more products, more technologically complex products, and more similar products. Conversely, countries which are less developed generally produce fewer, less complex, and less similar products.
Using econometric tests, I then determined that ECI Granger-causes GDP, while GDP does not Granger-cause ECI. This suggests that the economic situation of a country can actually be improved with intentional policy action. However, policymakers should focus on targeting coherent product clusters rather than raw export variety. Industrial policy and subsidies must prioritize shifting toward adjacent, highly complex sectors that share infrastructure and skill requirements with a country's existing competitive exports, minimizing transition costs while still growing the economy. Ultimately, moving up the value chain into complex, non-ubiquitous sectors secures a resilient path toward sustainable, long-run economic development.
Using econometric tests, I then determined that ECI Granger-causes GDP, while GDP does not Granger-cause ECI. This suggests that the economic situation of a country can actually be improved with intentional policy action. However, policymakers should focus on targeting coherent product clusters rather than raw export variety. Industrial policy and subsidies must prioritize shifting toward adjacent, highly complex sectors that share infrastructure and skill requirements with a country's existing competitive exports, minimizing transition costs while still growing the economy. Ultimately, moving up the value chain into complex, non-ubiquitous sectors secures a resilient path toward sustainable, long-run economic development.
Acknowledgements
Acknowledgements
I would like to express gratitude my mentor, Junseo Lee, for helping me immensely with so many aspects of this project. Additionally, I want to give thanks to Carlos Angulo and Amrita Kumar for their Economics help, to Rory, Eryn, Megan and Cyrus for their tireless work as directors of WSRP, and of course, to Stephen Wolfram.
Sources
Sources
◼
Dumitrescu, E.-I., & Hurlin, C. (2012). Testing for granger non-causality in heterogeneous panels. Economic Modelling, 29(4), 1450-1460. https://doi.org/10.1016/j.econmod.2012.02.014
◼
Economic Growth Rate [Dataset]. (n.d.). National Statistics, Republic of China (Taiwan). Retrieved July 8, 2026, from https://eng.stat.gov.tw/Point.aspx?sid=t.1&n=4200&sms=11713
◼
Emirmahmutoglu, F., & Kose, N. (2011). Testing for granger causality in heterogeneous mixed panels. Economic Modelling, 28(3), 870-876. https://doi.org/10.1016/j.econmod.2010.10.018
◼
Engle, R. F., & Granger, C. W. J. (1987). Co-Integration and error correction: Representation, estimation, and testing. Econometrica, 55(2), 251. https://doi.org/10.2307/1913236
◼
Hidalgo, C. A., & Hausmann, R. (2009). The building blocks of economic complexity. Proceedings of the National Academy of Sciences, 106(26), 10570-10575. https://doi.org/10.1073/pnas.0900943106
◼
Kejriwal, M., & Luo, Y. (2023). On the Empirical association between trade network complexity and Global gross domestic product. Studies in Computational Intelligence, 456-466. https://doi.org/10.1007/978-3-031-21127-0_37
◼
Koch, P. (2021). Economic complexity and growth: Can value-added exports better explain the link? Economics Letters, 198, 109682. https://doi.org/10.1016/j.econlet.2020.109682
◼
Mealy, P., Farmer, J. D., & Teytelboym, A. (2017). A new interpretation of the economic complexity index. SSRN Electronic Journal. https://doi.org/10.2139/ssrn.3075591
◼
Rao, C. (1982). Diversity and dissimilarity coefficients: A unified approach. Theoretical Population Biology, 21(1), 24-43. https://doi.org/10.1016/0040-5809(82)90004-1
◼
Toda, H. Y., & Yamamoto, T. (1995). Statistical inference in vector autoregressions with possibly integrated processes. Journal of Econometrics, 66(1-2), 225-250. https://doi.org/10.1016/0304-4076(94)01616-8
◼
Total Population [Dataset]. (n.d.). National Statistics, Republic of China (Taiwan). Retrieved July 8, 2026, from https://eng.stat.gov.tw/Point.aspx?sid=t.9&n=4208&sms=11713
◼
World Customs Organization. (1992). Harmonized commodity description and coding system (2nd ed.). https://www.wcoomd.org/
CITE THIS NOTEBOOK
CITE THIS NOTEBOOK
Evaluating the relationship between export diversification and economic development
by Jacob Chung
Wolfram Community, STAFF PICKS, July 9, 2026
https://community.wolfram.com/groups/-/m/t/3752366
by Jacob Chung
Wolfram Community, STAFF PICKS, July 9, 2026
https://community.wolfram.com/groups/-/m/t/3752366