In one of my earlier posts I wrote about installing Virtuoso triple store on Unix. In this post I thought I will put the links here for my future reference.
32-bit or 64-bit binary files are available here.
Installation steps are very well documented here. However do the Setup DSN only after creating a windows service for the default database.
Virtuoso is a fully functional database. It has in-built table called RDF_QUAD where it stores the triples.
If you have a scenario where in you want to convert the data that you have in your relational database into triples, you could use virtuoso to do it, however its not part of the open source release, but part of their commercial release.
Wednesday, March 10, 2010
Sunday, February 28, 2010
Cluster HeatMaps
I had heard of HeatMaps, but heard of Cluster HeatMaps very recently only when I had to generate one. It is being used a lot in Bioinformatics for gene expression studies.
Weinstein describes HeatMaps as following: (taken from here)
"In the case of gene expression data, the color assigned to a point in the heat map grid indicates how much of a particular RNA or protein is expressed in a given sample. The gene expression level is generally indicated by red for high expression and either green or blue for low expression. Coherent patterns (patches) of color are generated by hierarchical clustering on both horizontal and vertical axes to bring like together with like. Cluster relationships are indicated by tree-like structures adjacent to the heat map, and the patches of color may indicate functional relationships among genes and samples. "
I could not, for long time figure out how exactly they are generated. Reading here I got to understand that Cluster HeatMaps are generated by first performing hierarchial clustering for rows (i.e.genes) and then for columns(i.e. samples), generate their dendrograms and then use those for generating heat maps with trees for both genes and samples. I got to learn to do this in R(using hclust, heatmap functions) looking up here.(nice tutorial on R & Bioconductor)
Friday, February 12, 2010
Hierarchical Clustering in R
Clustering is very widely used technique in data analysis. It is classified into Unsupervised learning methos of data analysis in Machine Learning.
In the areas of gene expression data analysis, clustering is very helpful in getting to know about genes. If a gene belongs to a particular cluster and if we already know the functions of those genes present in the cluster, we can say that the gene also has functions as of those particular cluster of genes.
Most widely used clustering algorithms as K-means, Fuzzy-C-means, Hierarchial Clustering.
Hierarchial clustering can further be agglomerative or divisive. Look up here for a simple tutorial on the same.
I tried out Hierarchial clustering on PCR data in R.
I had a matric like this:
sample1 sample2 sample3 sample4 sample5
gene1
gene2
gene3
gene4
gene5
There are expression values (delta ct values)corresponding to each of the combination above. This data is read into R as a matrix usiing read.csv function.
To build hierarchial clusters, we need a Distance matrix or a similarity matrix.
For Distance matrix Euclidean Distance method can be used.
For Similarity matrix Pearson's distance can be used.
Pearson distance is given as d=1-r, where r is the Pearson's correlation coefficient.
In R dist() is the function to be used for euclidean distance.
distmatrix= dist(deltactmatrix,method="euclidean")
dist() doesn't have Pearson's method but there are other functions newly added, but I just did in the following manner.
simmatrix=as.dist(1-cor(deltactmatrix)).
Then use hr <- hclust(distmatrix/simmatrix, method="complete") to get clusters.
Then say plot(hr) to see the dendrograms.
Further steps include cutting the tree that is built at a point so that we can decide what clusters we want. There is a function called cut() specify the cluster and the height to cut at a point and then u get to see the clusters. Validating a cluster is also important and there are several methods (like bootstrapping hclusters & others) which one can find online in several papers.
In the areas of gene expression data analysis, clustering is very helpful in getting to know about genes. If a gene belongs to a particular cluster and if we already know the functions of those genes present in the cluster, we can say that the gene also has functions as of those particular cluster of genes.
Most widely used clustering algorithms as K-means, Fuzzy-C-means, Hierarchial Clustering.
Hierarchial clustering can further be agglomerative or divisive. Look up here for a simple tutorial on the same.
I tried out Hierarchial clustering on PCR data in R.
I had a matric like this:
sample1 sample2 sample3 sample4 sample5
gene1
gene2
gene3
gene4
gene5
There are expression values (delta ct values)corresponding to each of the combination above. This data is read into R as a matrix usiing read.csv function.
To build hierarchial clusters, we need a Distance matrix or a similarity matrix.
For Distance matrix Euclidean Distance method can be used.
For Similarity matrix Pearson's distance can be used.
Pearson distance is given as d=1-r, where r is the Pearson's correlation coefficient.
In R dist() is the function to be used for euclidean distance.
distmatrix= dist(deltactmatrix,method="euclidean")
dist() doesn't have Pearson's method but there are other functions newly added, but I just did in the following manner.
simmatrix=as.dist(1-cor(deltactmatrix)).
Then use hr <- hclust(distmatrix/simmatrix, method="complete") to get clusters.
Then say plot(hr) to see the dendrograms.
Further steps include cutting the tree that is built at a point so that we can decide what clusters we want. There is a function called cut() specify the cluster and the height to cut at a point and then u get to see the clusters. Validating a cluster is also important and there are several methods (like bootstrapping hclusters & others) which one can find online in several papers.
Wednesday, February 10, 2010
Given an InChI how to get its 3D co-ordinates
Recently, as part of the oreChem project that I am working for, I had to fetch 3D structure in CML format, given an InChI. Some of the InChI's I had were present in PubChem and some of them were not present in PubChem.
For those InChI's present in PubChem:
Its difficult to fetch the 3D structure given an InChI by doing a String search on the PubChem database.I found it time consuming even if there is an md5 Index on it. Easy way I figured out is to fetch the CID of that particular InChI, use that CID and then fetch the 3D structure.
Using eutils REST services appending the InChI in the the url here I got the CID of a given InChI. By Parsing the xml output and extracting the CID and using the following url here I got the 3D structure in XML format.
Another way is to convert InChI into SMILES using InChItoSMILES converter and then use smi23d to get 3D structure in SDF format, then convert SDF format into XML/CML using OpenEyeBabel or any other tool. (in Open Babel simply do "babel source.sdf result.cml")
For those not present in PubChem:
For these InChI's one could use OpenEyeBabel or OpenEyeChem 3D structure generating programs. I had posted this question on Blue Obelisk Stack exchange, a question and answer website for Cheminformaticians(similar to StackOverflow,SemanticOverflow). Here is the answer on how to do it. The 3D co-ordinated generated by OpenBabel and OpenEyeChem were not the same, they will not be same. Look here for the reasons.
If you need to Install Openbabel look here.
For those InChI's present in PubChem:
Its difficult to fetch the 3D structure given an InChI by doing a String search on the PubChem database.I found it time consuming even if there is an md5 Index on it. Easy way I figured out is to fetch the CID of that particular InChI, use that CID and then fetch the 3D structure.
Using eutils REST services appending the InChI in the the url here I got the CID of a given InChI. By Parsing the xml output and extracting the CID and using the following url here I got the 3D structure in XML format.
Another way is to convert InChI into SMILES using InChItoSMILES converter and then use smi23d to get 3D structure in SDF format, then convert SDF format into XML/CML using OpenEyeBabel or any other tool. (in Open Babel simply do "babel source.sdf result.cml")
For those not present in PubChem:
For these InChI's one could use OpenEyeBabel or OpenEyeChem 3D structure generating programs. I had posted this question on Blue Obelisk Stack exchange, a question and answer website for Cheminformaticians(similar to StackOverflow,SemanticOverflow). Here is the answer on how to do it. The 3D co-ordinated generated by OpenBabel and OpenEyeChem were not the same, they will not be same. Look here for the reasons.
If you need to Install Openbabel look here.
Wednesday, January 27, 2010
Schema less database: CouchDB
Thursday, December 3, 2009
Correct way of Shutting down Virtuoso server
Here is the correct way to shut down the virtuoso server.
At the SQL prompt type SHUTDOWN. This is the normal way of shutting down the virtuoso server.
For virtuoso in particular killing the virtuoso process would result in all transactions since the last checkpoint not being comitted to the database and remaining in the transaction log file, which would then have to be replayed when the server is next started. Whereas performing a normal shutdown will automatically run a checkpoint to commit all outstanding transactions to the database before shutting down.
For more information read at http://docs.openlinksw.com/virtuoso/isql.html#invokingisql
At the SQL prompt type SHUTDOWN. This is the normal way of shutting down the virtuoso server.
For virtuoso in particular killing the virtuoso process would result in all transactions since the last checkpoint not being comitted to the database and remaining in the transaction log file, which would then have to be replayed when the server is next started. Whereas performing a normal shutdown will automatically run a checkpoint to commit all outstanding transactions to the database before shutting down.
For more information read at http://docs.openlinksw.com/virtuoso/isql.html#invokingisql
Tuesday, November 17, 2009
Installing OpenLink Virtuoso triple store on Unix
Recently I had to install OpenLink virtuoso triple store for the orechem project that I am working for. They have an ocean of documentation. I was suggested by Marlon to put down the steps here for future reference. To learn what exactly Virtuoso is, read here.
This blog is about how I installed it on Unix. Will write down the windows installation in the next blog.
1) Downloaded the source files using wget
wget "http://sourceforge.net/projects/virtuoso/files/virtuoso/5.0.12/virtuoso-opensource-5.0.12.tar.gz/download"
2) untarred . tar -xvzf virtuoso-opensource-5.0.12.tar.gz
3) To generate the configure script it needs lot of other packages. Checked here for the list of package dependencies. Made sure they are installed.
4) cd into the directory created and typed ./configure. By default the install target directories are under /usr/local/ .In order to specify a particular target directory(in my case I created a directory called virtuoso-opensource) type ./configure --prefix pathtodir
5) Typed make (this took around 20-30 mins) followed by make install.
6) Four directories (var, bin, lib, share) were created in the configure target directory.
7) To start the virtuoso server configuration file "virtuoso.ini" is needed. I found that its located at var/lib/virtuoso/db
8) cd bin. Then started the server by typing:
./virtuoso-t -c ~/virtuoso-opensource/var/lib/virtuoso/db/virtuoso.ini -f
10) Accessed the web admin inerface from a web browser at http://gf18.ucs.indiana.edu:8890. I said gf18.ucs.indiana.edu, for I was working remotely on that machine. If thats not the case you can say localhost. 8890 is usually the port at which it is created. If this is not working check here for more information. The first time the virtuoso server starts, it installs Conductor VAD(Virtuoso Application Distribution) and an empty database.
11) On the web admin interface, clicked on Conductor and logged in using the default username dba and default password dba. There are lot of tabs there just explored those. From System Admin, User Accounts tab I created an account with user name schalla and set up a password for it.
12) Now comes the actual uploading of RDF triples. There are several methods described here. I implemented using the WebDAV browser available on the web interface and using "curl" from command line. To upload file from command line I typed :
curl -i -T bio2rdfdemo.rdf http://gf18.ucs.indiana.edu:8890/DAV/home/schalla/rdf_sink/bio2rdfdemo.rdf -u "schalla:password"
I could see the RDF files uploaded into rdf_sink folder. The actual process virtuoso follows is that it uploads RDF files into rdf_sink folder and from there the triples are stored on to the RDF_QUAD table in the database. To learn more about how exactly virtuoso stores triples read here.
13) Once RDF file is uploaded an IRI is to be generated. After a long search got the correct IRI generated, from RDF, Graphs tabs on the web interface. IRI created for the above file was http://local.virt/DAV/home/schalla/rdf_sink/bio2rdfdemo.rdf. Now using this as Default IRI graph in the SPARQL tab on web interface I executed the following sparql query
select DISTINCT ?p
where {?s ?p ?o}
and this query gave all the distinct properties in the graph.
Then tried running the SPARQL query from command line using curl as following
curl -F "query=SELECT DISTINCT ?p FROM WHERE {?s ?p ?o}" http://gf18.ucs.indiana.edu:8890/sparql
14) Port 1111 is the virtuoso DBMS port. This can be accesed using isql. First time when I typed ./isql 1111 dba dba I got an error that said "couldnotSQL connect"
Then checked to see the unixODBC drivers installed, so typed
odbcinst -j got this as output,
DRIVERS............: /etc/odbcinst.ini
SYSTEM DATA SOURCES: /etc/odbc.ini
USER DATA SOURCES..: /globalhome/schalla/.odbc.ini
After some reading on the web here and here got to learn that I need to include the following in the User Data sources i.e. .odbc.ini file.
[LocalVirt]
Driver=/globalhome/schalla/virtuoso-opensource/lib/virtodbc.so
Address=localhost:1111
[ODBC Data Sources]
triples-store=OpenLink Virtuoso
[triples-store]
Driver=OpenLink Virtuoso
Address=localhost:1111
Then again when I typed ./isql 1111 dba dba I got the SQL prompt and I could see the tables by typing "tables;" isql commands are not the same as SQL. Check here for more information.
15) How to terminate this virtuoso server ? I did not know how to do this. Thought there could be some command to do the same, searched a lot in the documentation with no luck, but what I finally did was killed the virtuoso process that was running.
Some of the useful links are here, here and here.
Virtuoso's Sponger RDFizer is a cartridge that can generate RDF data from non-RDF data (could be XML, HTML). I need to convert atom feed into RDF using GRDDL and I need to look into how I do it using sponger does this. Will post that stuff in the next blog.
This blog is about how I installed it on Unix. Will write down the windows installation in the next blog.
1) Downloaded the source files using wget
wget "http://sourceforge.net/projects/virtuoso/files/virtuoso/5.0.12/virtuoso-opensource-5.0.12.tar.gz/download"
2) untarred . tar -xvzf virtuoso-opensource-5.0.12.tar.gz
3) To generate the configure script it needs lot of other packages. Checked here for the list of package dependencies. Made sure they are installed.
4) cd into the directory created and typed ./configure. By default the install target directories are under /usr/local/ .In order to specify a particular target directory(in my case I created a directory called virtuoso-opensource) type ./configure --prefix pathtodir
5) Typed make (this took around 20-30 mins) followed by make install.
6) Four directories (var, bin, lib, share) were created in the configure target directory.
7) To start the virtuoso server configuration file "virtuoso.ini" is needed. I found that its located at var/lib/virtuoso/db
8) cd bin. Then started the server by typing:
./virtuoso-t -c ~/virtuoso-opensource/var/lib/virtuoso/db/virtuoso.ini -f
10) Accessed the web admin inerface from a web browser at http://gf18.ucs.indiana.edu:8890. I said gf18.ucs.indiana.edu, for I was working remotely on that machine. If thats not the case you can say localhost. 8890 is usually the port at which it is created. If this is not working check here for more information. The first time the virtuoso server starts, it installs Conductor VAD(Virtuoso Application Distribution) and an empty database.
11) On the web admin interface, clicked on Conductor and logged in using the default username dba and default password dba. There are lot of tabs there just explored those. From System Admin, User Accounts tab I created an account with user name schalla and set up a password for it.
12) Now comes the actual uploading of RDF triples. There are several methods described here. I implemented using the WebDAV browser available on the web interface and using "curl" from command line. To upload file from command line I typed :
curl -i -T bio2rdfdemo.rdf http://gf18.ucs.indiana.edu:8890/DAV/home/schalla/rdf_sink/bio2rdfdemo.rdf -u "schalla:password"
I could see the RDF files uploaded into rdf_sink folder. The actual process virtuoso follows is that it uploads RDF files into rdf_sink folder and from there the triples are stored on to the RDF_QUAD table in the database. To learn more about how exactly virtuoso stores triples read here.
13) Once RDF file is uploaded an IRI is to be generated. After a long search got the correct IRI generated, from RDF, Graphs tabs on the web interface. IRI created for the above file was http://local.virt/DAV/home/schalla/rdf_sink/bio2rdfdemo.rdf. Now using this as Default IRI graph in the SPARQL tab on web interface I executed the following sparql query
select DISTINCT ?p
where {?s ?p ?o}
and this query gave all the distinct properties in the graph.
Then tried running the SPARQL query from command line using curl as following
curl -F "query=SELECT DISTINCT ?p FROM
14) Port 1111 is the virtuoso DBMS port. This can be accesed using isql. First time when I typed ./isql 1111 dba dba I got an error that said "couldnotSQL connect"
Then checked to see the unixODBC drivers installed, so typed
odbcinst -j got this as output,
DRIVERS............: /etc/odbcinst.ini
SYSTEM DATA SOURCES: /etc/odbc.ini
USER DATA SOURCES..: /globalhome/schalla/.odbc.ini
After some reading on the web here and here got to learn that I need to include the following in the User Data sources i.e. .odbc.ini file.
[LocalVirt]
Driver=/globalhome/schalla/virtuoso-opensource/lib/virtodbc.so
Address=localhost:1111
[ODBC Data Sources]
triples-store=OpenLink Virtuoso
[triples-store]
Driver=OpenLink Virtuoso
Address=localhost:1111
Then again when I typed ./isql 1111 dba dba I got the SQL prompt and I could see the tables by typing "tables;" isql commands are not the same as SQL. Check here for more information.
15) How to terminate this virtuoso server ? I did not know how to do this. Thought there could be some command to do the same, searched a lot in the documentation with no luck, but what I finally did was killed the virtuoso process that was running.
Some of the useful links are here, here and here.
Virtuoso's Sponger RDFizer is a cartridge that can generate RDF data from non-RDF data (could be XML, HTML). I need to convert atom feed into RDF using GRDDL and I need to look into how I do it using sponger does this. Will post that stuff in the next blog.
Subscribe to:
Posts (Atom)