Based on health industry interviews and surveys with over 2,000
individuals, ON World’s recently published report analyzes the market
for mobile sensing solutions for health and wellness.
San Diego, CA, May 30, 2013-- Mobile sensing technologies
are revolutionizing healthcare by providing low-cost patient-centered
solutions with real-time patient to doctor monitoring, according to ON
World, a global technology research firm.
“Advances in low power wireless communications, MEMS and
multi-sensor arrays have resulted in viable body area network
applications for clinical patient monitoring, assisted care, at-home
chronic disease management and general wellness,” says Mareca Hatler,
ON World’s research director. “At the same time, there has been
enormous growth for mobile sensing apps for smartphones and tablets.”
By 2017, there will be 1.4 billion mobile sensing health and
fitness app downloads worldwide and health apps will increase the
fastest over the next five years.
Mobile sensing solutions include wearable and implantable
sensors as well as carry-able devices that can be used while the user
is mobile. Global annual sensor shipments for mobile sensing health
and fitness devices-- including dedicated devices and health/fitness
enabled smart devices such as smart watches, smartphones and tablets—
will reach 515 million in 2017 up from 107 million in 2012.
ON World’s recent survey (Q2:2013) with over 1,000 U.S.
consumers found that nearly 4 in 10 are interested in purchasing a
smart watch and 48% are likely to use their smart watch for health or
fitness. Respondents are especially interested in using their smart
watch for blood pressure and heart rate monitoring as well as activity
tracking.
Traditional health products are quickly being replaced with
mobile sensing solutions. The pace of innovation in this area is
illustrated by the fact that 60% of the mobile sensing products ON World
evaluated were launched in 2012 or later. Cardiac/ECG monitoring is a
popular developer area that makes up 22% of the health products ON
World researched. There are a growing number of innovations targeting
diabetes management such as Cellnovo’s cellular connected continuous
blood glucose monitoring system, smartphone/cellular connected blood
glucose meters by BodyTel, Entra Health, Sanofi/AgaMatrix and Telcare.
Breakthrough implantable sensor technology will be one of the fastest
growing markets where blood glucose sensors with multi-month lifetimes
will change the lives of diabetics and lower overall healthcare costs.
Although it is widespread for fitness, wearable multi-sensor
vital sign monitors is still an emerging market for healthcare. The
results from pilots so far confirm that wearable vital sign monitors
are effective and preferred by both patients and health workers.
Between 2012 and 2017, wearable health and fitness device shipments
will increase by 552% and make up over 80% of the mobile sensing health
and fitness device market at this time.
Based on input from more than 2,000 individuals, “Mobile Sensing Health & Wellness”
analyzes the global market for mobile sensing solutions for health and
wellness including 5-year forecasts for dedicated devices, passive and
active sensors, communication chipsets, health/fitness-enabled smart
mobile devices, smart watches, monitoring services, mobile apps and
subscriptions.
The report also includes primary research findings from
several recently completed surveys and an in-depth analysis of the
leading vendors. The technologies covered in the report include
Bluetooth Classic, Bluetooth Smart, WiFi, ZigBee, ANT, BodyLAN, NFC and
others. The report may be purchased separately or as part of a set
with another report, “Mobile Sensing Sports & Fitness.”
The report synopsis and a free executive summary is available from: http://www.onworld.com/mobilesensing/health
Nguyen The Vinh
Đại Học Thái Nguyên
Thứ Ba, 17 tháng 9, 2013
95 Million Bluetooth Chipset Shipments for Health & Fitness in 2017
ON World’s recent research covers rapidly growing mobile
sensing markets for health, wellness, sports and fitness including
wearable, implantable and other Body Area Network variations.
San Diego, CA, June 12, 2013-- Adoption of mobile sensing health and fitness products will increase by 500 percent over the next five years, says ON World, a global technology market research firm.
“Smartphones with Bluetooth Smart onboard provides mobile sensing developers with iOS and Android libraries, app stores, cloud connected mobile gateways and hundreds of millions of users to target their creations to,” says Mareca Hatler, ON World’s research director. “Bluetooth Smart delivers what mobile health and wellness solutions require: fast connections, ultra-low power consumption, low cost and interoperability.”
By 2017, Bluetooth chipsets used for health, wellness, sports and fitness will reach 95.7 million and much of the growth will be attributed to Bluetooth Smart used in mobile sensing devices. At this time, mobile sensing devices will make up 80% of the wireless health and fitness device shipments.
The migration to Bluetooth Smart for health and fitness is accelerating. Bluetooth Smart certified products have increased by almost 10 times over the past year. ON World’s evaluation of 200 mobile sensing health and fitness products found that Bluetooth Smart will be used in 59% of the products launched in 2013.
A few examples of some of the most recent and planned Bluetooth Smart health and fitness products include heart rate monitors by BeetsBlu, Runtastic, Polar and Scosche; activity trackers and sports watches by BodyMedia, Fitbit, Fitbug, Garmin and Withings; and vital sign monitors by Nonin, iHealth Lab and iMPak Health.
Based on input from more than 2,000 individuals, ON World’s recently published reports “Mobile Sensing Health & Wellness” and “Mobile Sensing Sports & Fitness” analyze the global market for mobile sensing solutions for healthcare, general wellness and sports/fitness. Each report provides 5-year forecasts for dedicated health/fitness devices, sensors, communication chipsets, health/fitness-enabled smart mobile devices, smart watches, monitoring services, mobile apps and subscriptions. The reports also contain the results from several recently completed surveys and an in-depth analysis of 100+ companies and their products/services. The technologies covered include Bluetooth Classic, Bluetooth Smart, WiFi, ZigBee, ANT, BodyLAN, NFC and others.
Each report may be purchased separately or as part of a set.
The report synopses and executive summaries for both reports are available from:
http://onworld.com/mobilesensing/health-fitness-set
San Diego, CA, June 12, 2013-- Adoption of mobile sensing health and fitness products will increase by 500 percent over the next five years, says ON World, a global technology market research firm.
“Smartphones with Bluetooth Smart onboard provides mobile sensing developers with iOS and Android libraries, app stores, cloud connected mobile gateways and hundreds of millions of users to target their creations to,” says Mareca Hatler, ON World’s research director. “Bluetooth Smart delivers what mobile health and wellness solutions require: fast connections, ultra-low power consumption, low cost and interoperability.”
By 2017, Bluetooth chipsets used for health, wellness, sports and fitness will reach 95.7 million and much of the growth will be attributed to Bluetooth Smart used in mobile sensing devices. At this time, mobile sensing devices will make up 80% of the wireless health and fitness device shipments.
The migration to Bluetooth Smart for health and fitness is accelerating. Bluetooth Smart certified products have increased by almost 10 times over the past year. ON World’s evaluation of 200 mobile sensing health and fitness products found that Bluetooth Smart will be used in 59% of the products launched in 2013.
A few examples of some of the most recent and planned Bluetooth Smart health and fitness products include heart rate monitors by BeetsBlu, Runtastic, Polar and Scosche; activity trackers and sports watches by BodyMedia, Fitbit, Fitbug, Garmin and Withings; and vital sign monitors by Nonin, iHealth Lab and iMPak Health.
Based on input from more than 2,000 individuals, ON World’s recently published reports “Mobile Sensing Health & Wellness” and “Mobile Sensing Sports & Fitness” analyze the global market for mobile sensing solutions for healthcare, general wellness and sports/fitness. Each report provides 5-year forecasts for dedicated health/fitness devices, sensors, communication chipsets, health/fitness-enabled smart mobile devices, smart watches, monitoring services, mobile apps and subscriptions. The reports also contain the results from several recently completed surveys and an in-depth analysis of 100+ companies and their products/services. The technologies covered include Bluetooth Classic, Bluetooth Smart, WiFi, ZigBee, ANT, BodyLAN, NFC and others.
Each report may be purchased separately or as part of a set.
The report synopses and executive summaries for both reports are available from:
http://onworld.com/mobilesensing/health-fitness-set
Thứ Sáu, 6 tháng 9, 2013
Preparing an effective job-seeking cover letter…
(Note – you might also want to check out my related post of tips for preparing your resume.)
I often get asked to provide advice and
feedback for engineering students (BSc, MSc and PhD) applying for jobs
and though the application contents can vary widely, especially between
industry and academia, one thing that all employers have in common is
the requirement for a cover letter. It’s probably the most important
part of the application package, yet it seems to be the part most people
do poorly. In this post, I’ll try to help you out by presenting some of
the tips I generally suggest to my own students based on: a) what I’ve
learned from others, b) what has worked for me when I’ve applied for
jobs, and c) what I like to see when I am looking to hire someone.
The single most important objective of
the cover letter is to convince the company, agency or university that
you are applying to that they cannot afford to pass up the opportunity
to hire you. You do this by illustrating what a tremendous asset you
will be to them. In that context – they don’t care what sort of career
you want or what experience they can offer you – they want to hear what
you can offer them. The time to see what they can do for you is later on
– once they realize they cannot get by without you and you begin
negotiating the details of your employment. Here below are my top ten
tips for writing a compelling cover letter…
-
Research the company, agency, or university. Find out as much as you can about the people you will be working with and the type of work they do. Keep this knowledge in mind as you write the letter and try to find ways to illustrate that you’ve done this. There is usually lots of info on the internet and companies expect you to research them before applying. In fact, they’re usually insulted when you don’t.
-
Provide your name, mailing address, phone number and email address in a letterhead or return address header.
-
Use a Subject Line and cite a specific position title and number if available.
-
Address your letter to a specific person, if possible. Employ a proper salutation (i.e. “Dear Ms. Jones”) and do not use first names, even if you know the person. If you don’t have a specific person to address, open with “Dear Sir or Madam”. Never assume the gender of the reader – I can’t tell you how many grad student applications I ignore each year because they’ve addressed the cover letter to “Dear Sir” – probably hundreds. (This goes back to Hint # 1.)
-
Use the cover letter to explain and highlight things in your resume that are relevant to them. That is, customize your cover letter to their particular company and/or job advertisement. (The really sharp applicant actually customizes their resume to the particular job, as well!)
-
Remember, employers use cover letters to assess your writing skills and your attention to detail. So write in proper paragraphs (i.e. with topic sentences and supporting facts). Don’t write in point form – it makes it look like a resume not a letter. Avoid abbreviations and acronyms (especially undefined ones) and don’t use casual phrases, slang or an overly familiar tone. Finally, proofread the letter carefully; make sure there are no spelling or grammar errors.
-
Try to keep it to about 1 page. You can push things a bit by using Times 11 and 0.75 inch borders, or you can go up to about a page and a quarter – but two full pages is too much. People usually have a lot to read and two pages of tightly packed prose is a real put-off. Also, don’t use fonts smaller than Times 11 – most of the people in charge (i.e. the ones who decide on the hiring) are old enough to need reading glasses and small fonts are extremely frustrating to them.
-
If you have experience or preferences to do a particular type of work, don’t list them as the type of projects you want, or expect, to work on. Instead, give them as examples of the types of projects you could take on right away with minimal guidance and supervision. (Again, it’s all about what you can do for them, not vice versa.) Also, don’t suggest things they don’t do… (That’s one sure way to demonstrate that you’ve ignored Hint #1.)
-
Highlight your communications skills. Have you written any reports or papers? Have you presented papers or posters at conferences? Give specific examples of your oral and written communications skills.
-
Close the letter by stating specifically when you could start work and when exactly you are available for an interview. Ask explicitly for the interview. Don’t forget a proper closing salutation (“Yours truly,” or “Sincerely,” are the most common and appropriate) and type your name in below where you will sign.
Once you’ve written the draft – leave it
for a day, then go back and give it a critical inspection. Are you
effectively presenting the “you attitude” instead of the “me attitude”?
Are you effectively demonstrating your technical writing skills (i.e.
proper paragraphing) and meticulous attention to detail (e.g. spelling
and grammar)? Have you effectively illustrated you suitability for this
particular job? It’s always a good idea to get a friend or mentor (e.g.
your professor) to read the letter and provide feedback.
Finally SIGN YOU LETTER! Yes, get an actual pen and sign the letter!
Good luck! Let me know if you get the job!
How to write to a prospective PhD (or Post-Doc) supervisor
I’ve noticed recently that a lot of
people make their way to this site while searching for advice on how to
write to a potential PhD supervisor. I’ve also noticed that many of the
letters/emails that I personally get on this topic are actually
irrelevant to me, poorly written, or both. So, I figured it might be a
good idea to put a little bit of advice out there to help students who
are trying to get into a PhD (or Masters) program. All the same
principles also apply for those seeking post-doc supervisors.
There are two categories of people
writing these letters – those who need financial support for
their graduate program and those who don’t. If you have a big
scholarship, or if you’re independently wealthy, you fall into the
second category and it’s important to mention it right at the beginning
of your letter. Believe me – that will get you noticed! The reason is
that the other 99% of people are looking for a supervisor AND for
financial support, and there’s just not enough money to go around.
Letting a prospective supervisor know that you are not looking for money
increases the likelihood that they will read the rest of your letter
and, as a result, it improves your chances of getting admitted.
You might be wondering, “Do I even need to write letters to potential supervisors? Can’t I just fill out university application forms?” The
fact is – hundreds of people apply for each spot and if you just fill
out the application form without actually contacting any of the
professors at the university in question, then you’re not likely to get
noticed. Furthermore, most graduate students (especially PhDs) and
Post-Docs are admitted because a particular professor has expressed an
interest in recruiting them, so if you’ve not been in contact with a
professor, chances are, nobody there is going to push to get you
admitted.
The next obvious question is whether you
have to write an actual letter, or will an email do? The answer is
‘both’. You should do it by email, but write it in the form of a proper
letter. Specifically: make the subject line informative, use a proper
salutation, write in proper paragraphs,
organize your thoughts, make sure the spelling and grammar are perfect,
and end it with a proper closing. Essentially, you are applying for a
job and so, in a way, your application form is analogous to your resume
and your letter to a prospective supervisor is equivalent to the
corresponding cover letter.
Like those who write a good cover letter when applying for a job,
students who write good letters to potential supervisors are more likely
to get noticed.
You can go ahead and read about writing an effective cover letter
to get some basic advice on witting to a potential PhD (or Post-doc, or
Masters) supervisor. Here below are some more specific tips for you.
(You’ll notice a bit of overlap.)
Do Your Research
It’s important to write to a specific
person about doing a specific type of research. I get all sorts of
emails addressed to ‘Dear Sir’ (with 20 other people in the address
line). They all go right in my trash folder. In the first
place, anyone who actually thinks that it is acceptable to assume that
all professors are men is living in the 19th century and is probably
totally out of touch with the current literature and technology, as
well. In the second place, I assume (as does every other prof out there)
that if 19 other people got the same email, then you’ll be just as
happy if one of them answers your email - so I’m not going to waste my
time on it. Furthermore, I study river ice – nothing else – I don’t
plan to do any projects on groundwater, water resources planning and
management, construction, chemical engineering or nuclear physics. Yet I
get tons of people emailing me, asking if they can come and do research
with me on these (and many other completely irrelevant) topics, and
they actually expect me to pay them to do it! I delete all of these
emails, too. If you’re doing this sort of thing in the blind hope that
you might get lucky and hit on just one professor whose interests
intersect with yours, you are wasting your time completely. Think about
it - you’re trying to land a research position and you haven’t even
bothered to do the most trivial research on the topic (i.e. surf the web
and actually find out who is doing research that matches your interests
and experience). Every professor that reads your email is going to
think that you are either totally lazy or completely inept as a
researcher (probably both). They’re definitely not going to have any
interest in recruiting you.
Your best chance at getting someone
enthused about recruiting you is to find someone whose interests match
your own. Therefore, as an absolute minimum, you should check out their
website to see if they do anything even remotely related to your
area(s) of interest. If they don’t, then you’re just wasting your time
(and theirs) by writing to them. It’s also important to keep in mind
that all professors have well-defined research programs and they seek
out and get money to support those specific research programs. So, it’s
important to demonstrate an interest in their research projects, not to
simply dictate your own research interests to them.
It’s true that many people don’t have a
specific PhD (or Masters) topic in mind and will accept almost any
project just to get an opportunity to do a PhD. That’s perfectly fine;
go ahead and admit it. In fact, I encourage you to write to prospective
professors and ask them what they are working on and whether they might
have any projects for which they are seeking graduate students.
Personally, I am much more inclined to follow-up with an applicant who
does this, than with one who tells me what they plan to work on,
especially when it’s irrelevant to me.
Demonstrate Your Relevant Merits
Here again, it’s important to do your
research. There is no point in applying to a graduate program if you
don’t have the grades to get in, yet a surprising number of people do.
Most universities post their academic requirements on their web sites –
check them out and keep in mind, these are minimum requirements.
Meeting these minimum requirements will not necessarily get you
admitted, especially if you don’t have a specific professor asking for
you. You should be aware that academic requirements may also vary by
program. For example, at most Canadian universities, you need a
Master’s degree to get admitted to an engineering PhD program, whereas
that might not be the case for PhD programs in science. In your letter
to the prospective supervisor, make it clear that you have checked these
academic requirements and that you exceed them all. If you have done a
Masters, be sure to mention the title of your thesis and the name of
your thesis supervisor in your letter.
You also need to demonstrate that your
academic background is relevant to that professor’s research program.
For example, as a hydrotechnical engineer who specializes in river ice, I
am not likely to recruit someone who did a Masters in environmental
engineering. We may both be civil engineers, but that’s not a
particularly relevant background for a PhD in
hydrotechnical engineering. In fact, relevant skills can be as specific
as the type of research experience you have. In this context, you
really should download and read a few of the professors’ journal papers
to get an idea of what types of expertise they might be seeking. For
example, if someone is doing numerical modeling and you have experience
in that (even if only a single graduate course) then be sure to mention
it in your letter. You’re far more likely to spark their interest than
someone who has absolutely no experience or expertise in modeling. The
research experience expectations tend to be quite a bit less
rigourous for prospective Masters students. A relevant undergraduate
degree is typically essential and any sort of research experience (e.g. a
summer or co-op research job) is an asset but often not essential.
Doing research on the professors that
you’ll be contacting not only ensures you’ll be approaching the
appropriate people, it will increase your chances of attracting their
interest, since it’s a very real demonstration of your initiative,
curiosity and resourcefulness.
Keep it Brief
Many of the letters I get from
prospective PhD students are excessively long (i.e. a couple of pages or
more). If I open an email to find such a long letter, I usually close
it for the moment, with the intention of looking at it later when I have
more time. However, I receive over 100 emails a day, so it’s usually
forgotten by the next day. Sometimes I mark them for follow-up, but
I’ve got about a hundred emails flagged at any given time – so it still
might get lost in the shuffle. In contrast, if an email is only one or
two (real) paragraphs long, I read it right away. I’m not unique in
this; in fact, many professors get several hundred emails a day and read
only a few of them. Keep that in mind as you write your letter and
make a concerted effort to be brief. Aim to get your message across in
two paragraphs at the most. The goal is to spark the professor’s
interest in order to initiate a dialog; you don’t need to tell them your
whole life’s story in the first contact.
Put Something Meaningful in the Subject Line
Most people who receive excessive amounts
of email, like professors, prioritize what they read based on the
subject lines. If your subject line is blank, simply says “c.v.”, or
even worse says “hey professor” – it may be ignored indefinitely. For
obvious reasons, the subject line that would catch my eye immediately is
“Prospective PhD student seeking to study river ice”. In 23 years as a professor I never received a single email with this subject line (until I wrote this post
).
Attach Supporting Material
Scan copies of your transcripts and
attach them to the email along with a copy of your resume. You’ll have
to send official paperwork for the application process, but it takes
time for a professor to go hunt that up. If you save them that time by
providing the info for them, they’re more likely to follow-up. Remember
though, all unofficial transcripts are eventually compared against the
official versions. Any discrepancies, no matter how minor, are
guaranteed to kill any chance of admission.
If you have published any journal or
conference papers, include them as attachments to your email. This not
only shows evidence of your research productivity, it gives the
professor a better idea of your research background and some indication
of your writing skills.
Do you need to do this to get into a Masters program?
What if you’re an undergraduate seeking
admission to a Masters program? Should you write letters to prospective
MSc supervisors? That depends upon whether there is a particular topic
you’d like to study. If yes, then it makes sense to contact professors
working in that research area to see if they are willing to take you
on.
Good luck !
Source: http://thesistips.wordpress.com
Thứ Ba, 3 tháng 9, 2013
Tôi chọn Java, Hive hoặc Pig như thế nào?
Bạn có nhiều tùy chọn để lập trình Hadoop và tốt nhất là xem xét trường hợp sử dụng để chọn đúng công cụ cho công việc đó. Bạn không bị hạn chế chỉ làm việc với dữ liệu quan hệ nhưng bài này tập trung vào Informix, DB2 và Hadoop để hoạt động tốt với chúng. Việc viết hàng trăm dòng mã bằng Java để thực hiện một phép nối băm kiểu quan hệ là hoàn toàn phí thời gian do thuật toán MapReduce của Hadoop đã có sẵn. Bạn chọn tùy chọn nào? Đó là vấn đề sở thích cá nhân. Một số người thích viết mã các phép toán tập hợp bằng SQL. Một số người khác thích viết mã kiểu thủ tục. Bạn nên chọn ngôn ngữ mà bạn làm việc hiệu quả nhất. Nếu bạn có nhiều hệ thống quan hệ và muốn kết hợp tất cả các dữ liệu với hiệu năng lớn với một mức giá thấp, thì Hadoop, MapReduce, Hive và Pig đã sẵn sàng để trợ giúp.
Hầu hết các cơ sở dữ liệu quan hệ hiện đại đều có thể phân vùng dữ liệu. Một trường hợp sử dụng phổ biến là phân vùng theo khoảng thời gian. Một cửa sổ thời gian cố định của các dữ liệu được lưu trong cơ sở dữ liệu, ví dụ một khoảng thời gian 18 tháng trôi qua, sau đó dữ liệu được đưa vào lưu trữ. Khả năng tách phân vùng là rất mạnh. Nhưng sau khi phân vùng được tách ra người ta làm gì với dữ liệu?
Việc lưu trữ dữ liệu cũ vào băng từ là một cách rất tốn kém để loại bỏ các byte dữ liệu cũ. Sau khi đã di chuyển sang môi trường ít có thể truy cập hơn, dữ liệu rất hiếm khi được truy cập trừ khi có một yêu cầu kiểm toán hợp pháp. Hadoop đưa ra một sự thay thế khác tốt hơn.
Di chuyển các byte dữ liệu lưu trữ từ các phân vùng cũ vào Hadoop tạo ra khả năng truy cập hiệu năng cao với chi phí thấp hơn nhiều so với việc duy trì dữ liệu trong hệ thống giao dịch hoặc quầy dữ liệu/kho dữ liệu ban đầu. Dữ liệu là quá cũ không còn giá trị giao dịch nữa, nhưng vẫn còn rất có giá trị với tổ chức để dùng cho các phân tích dài hạn. Các ví dụ Sqoop được hiển thị ở trên cung cấp những điều cơ bản về cách di chuyển dữ liệu này từ một phân vùng quan hệ sang HDFS.
Có thể truy cập dữ liệu Informix/DB2/tệp phẳng trong HDFS thông qua NFS, như hiển thị trong Liệt kê 20. Cách này cung cấp các hoạt động dòng lệnh mà không cần sử dụng giao diện "hadoop fs-yadayada". Theo quan điểm công nghệ về trường hợp sử dụng, NFS rất bị hạn chế trong một môi trường Dữ liệu lớn, nhưng các ví dụ được cung cấp cho các nhà phát triển và dữ liệu không-lớn-lắm.
Liệt kê 20. Thiết lập Fuse - truy cập dữ liệu HDFS của bạn thông qua NFS
# this is for CDH4, the CDH3 image doesn't have fuse installed...
$ mkdir fusemnt
$ sudo hadoop-fuse-dfs dfs://localhost:8020 fusemnt/
INFO fuse_options.c:162 Adding FUSE arg fusemnt/
$ ls fusemnt
tmp user var
$ ls fusemnt/user
cloudera hive
$ ls fusemnt/user/cloudera
customer DS.txt.gz HF.out HF.txt orders staff
$ cat fusemnt/user/cloudera/orders/part-m-00001
1007,2008-05-31,117,null,n,278693 ,2008-06-05,125.90,25.20,null
1008,2008-06-07,110,closed Monday
,y,LZ230 ,2008-07-06,45.60,13.80,2008-07-21
1009,2008-06-14,111,next door to grocery
,n,4745 ,2008-06-21,20.40,10.00,2008-08-21
1010,2008-06-17,115,deliver 776 King St. if no answer
,n,429Q ,2008-06-29,40.60,12.30,2008-08-22
1011,2008-06-18,104,express
,n,B77897 ,2008-07-03,10.40,5.00,2008-08-29
1012,2008-06-18,117,null,n,278701 ,2008-06-29,70.80,14.20,null
|
Thế hệ tiếp theo của Flume hay là flume-ng là trình nạp song song tốc độ cao. Các cơ sở dữ liệu có các trình nạp tốc độ cao, vậy làm thế nào để chúng sẽ hoạt động tốt với nhau? Trường hợp sử dụng dữ liệu quan hệ dành cho Flume-ng sẽ tạo ra một tệp sẵn sàng để nạp, tại chỗ hoặc từ xa, sao cho một máy chủ dữ liệu quan hệ có thể dùng trình nạp tốc độ cao của mình. Đúng là chức năng này chồng lên Sqoop, nhưng kịch bản lệnh được hiển thị trong Liệt kê 21 đã được tạo ra theo yêu cầu của một khách hàng đặc biệt cho kiểu nạp cơ sở dữ liệu này.
Liệt kê 21. Xuất khẩu dữ liệu HDFS tới một tệp phẳng để nạp bởi một cơ sở dữ liệu
$ sudo yum install flume-ng
$ cat flumeconf/hdfs2dbloadfile.conf
#
# started with example from flume-ng documentation
# modified to do hdfs source to file sink
#
# Define a memory channel called ch1 on agent1
agent1.channels.ch1.type = memory
# Define an exec source called exec-source1 on agent1 and tell it
# to bind to 0.0.0.0:31313. Connect it to channel ch1.
agent1.sources.exec-source1.channels = ch1
agent1.sources.exec-source1.type = exec
agent1.sources.exec-source1.command =hadoop fs -cat /user/cloudera/orders/part-m-00001
# this also works for all the files in the hdfs directory
# agent1.sources.exec-source1.command =hadoop fs
# -cat /user/cloudera/tsortin/*
agent1.sources.exec-source1.bind = 0.0.0.0
agent1.sources.exec-source1.port = 31313
# Define a logger sink that simply file rolls
# and connect it to the other end of the same channel.
agent1.sinks.fileroll-sink1.channel = ch1
agent1.sinks.fileroll-sink1.type = FILE_ROLL
agent1.sinks.fileroll-sink1.sink.directory =/tmp
# Finally, now that we've defined all of our components, tell
# agent1 which ones we want to activate.
agent1.channels = ch1
agent1.sources = exec-source1
agent1.sinks = fileroll-sink1
# now time to run the script
$ flume-ng agent --conf ./flumeconf/ -f ./flumeconf/hdfs2dbloadfile.conf -n
agent1
# here is the output file
# don't forget to stop flume - it will keep polling by default and generate
# more files
$ cat /tmp/1344780561160-1
1007,2008-05-31,117,null,n,278693 ,2008-06-05,125.90,25.20,null
1008,2008-06-07,110,closed Monday ,y,LZ230 ,2008-07-06,45.60,13.80,2008-07-21
1009,2008-06-14,111,next door to ,n,4745 ,2008-06-21,20.40,10.00,2008-08-21
1010,2008-06-17,115,deliver 776 King St. if no answer ,n,429Q
,2008-06-29,40.60,12.30,2008-08-22
1011,2008-06-18,104,express ,n,B77897 ,2008-07-03,10.40,5.00,2008-08-29
1012,2008-06-18,117,null,n,278701 ,2008-06-29,70.80,14.20,null
# jump over to dbaccess and use the greatest
# data loader in informix: the external table
# external tables were actually developed for
# informix XPS back in the 1996 timeframe
# and are now available in may servers
#
drop table eorders;
create external table eorders
(on char(10),
mydate char(18),
foo char(18),
bar char(18),
f4 char(18),
f5 char(18),
f6 char(18),
f7 char(18),
f8 char(18),
f9 char(18)
)
using (datafiles ("disk:/tmp/myfoo" ) , delimiter ",");
select * from eorders;
|
Oozie sẽ xâu chuỗi nhiều tác vụ của Hadoop với nhau. Có một tập hợp các ví dụ hấp dẫn kèm theo oozie, đã được dùng trong đoạn mã được hiển thị trong Liệt kê 22.
Liệt kê 22. Kiểm soát tác vụ bằng oozie
# This sample is for CDH3
# untar the examples
# CDH4
$ tar -zxvf /usr/share/doc/oozie-3.1.3+154/oozie-examples.tar.gz
# CDH3
$ tar -zxvf /usr/share/doc/oozie-2.3.2+27.19/oozie-examples.tar.gz
# cd to the directory where the examples live
# you MUST put these jobs into the hdfs store to run them
$ hadoop fs -put examples examples
# start up the oozie server - you need to be the oozie user
# since the oozie user is a non-login id use the following su trick
# CDH4
$ sudo su - oozie -s /usr/lib/oozie/bin/oozie-sys.sh start
# CDH3
$ sudo su - oozie -s /usr/lib/oozie/bin/oozie-start.sh
# checkthe status
oozie admin -oozie http://localhost:11000/oozie -status
System mode: NORMAL
# some jar housekeeping so oozie can find what it needs
$ cp /usr/lib/sqoop/sqoop-1.3.0-cdh3u4.jar examples/apps/sqoop/lib/
$ cp /home/cloudera/Informix_JDBC_Driver/lib/ifxjdbc.jar examples/apps/sqoop/lib/
$ cp /home/cloudera/Informix_JDBC_Driver/lib/ifxjdbcx.jar examples/apps/sqoop/lib/
# edit the workflow.xml file to use your relational database:
#################################
<command> import
--driver com.informix.jdbc.IfxDriver
--connect jdbc:informix-sqli://192.168.1.143:54321/stores_demo:informixserver=ifx117
--table orders --username informix --password useyours
--target-dir /user/${wf:user()}/${examplesRoot}/output-data/sqoop --verbose<command>
#################################
# from the directory where you un-tarred the examples file do the following:
$ hrmr examples;hput examples examples
# now you can run your sqoop job by submitting it to oozie
$ oozie job -oozie http://localhost:11000/oozie -config \
examples/apps/sqoop/job.properties -run
job: 0000000-120812115858174-oozie-oozi-W
# get the job status from the oozie server
$ oozie job -oozie http://localhost:11000/oozie -info 0000000-120812115858174-oozie-oozi-W
Job ID : 0000000-120812115858174-oozie-oozi-W
-----------------------------------------------------------------------
Workflow Name : sqoop-wf
App Path : hdfs://localhost:8020/user/cloudera/examples/apps/sqoop/workflow.xml
Status : SUCCEEDED
Run : 0
User : cloudera
Group : users
Created : 2012-08-12 16:05
Started : 2012-08-12 16:05
Last Modified : 2012-08-12 16:05
Ended : 2012-08-12 16:05
Actions
----------------------------------------------------------------------
ID Status Ext ID Ext Status Err Code
---------------------------------------------------------------------
0000000-120812115858174-oozie-oozi-W@sqoop-node OK
job_201208120930_0005 SUCCEEDED -
--------------------------------------------------------------------
# how to kill a job may come in useful at some point
oozie job -oozie http://localhost:11000/oozie -kill
0000013-120812115858174-oozie-oozi-W
# job output will be in the file tree
$ hcat /user/cloudera/examples/output-data/sqoop/part-m-00003
1018,2008-07-10,121,SW corner of Biltmore Mall ,n,S22942
,2008-07-13,70.50,20.00,2008-08-06
1019,2008-07-11,122,closed till noon Mondays ,n,Z55709
,2008-07-16,90.00,23.00,2008-08-06
1020,2008-07-11,123,express ,n,W2286
,2008-07-16,14.00,8.50,2008-09-20
1021,2008-07-23,124,ask for Elaine ,n,C3288
,2008-07-25,40.00,12.00,2008-08-22
1022,2008-07-24,126,express ,n,W9925
,2008-07-30,15.00,13.00,2008-09-02
1023,2008-07-24,127,no deliveries after 3 p.m. ,n,KF2961
,2008-07-30,60.00,18.00,2008-08-22
# if you run into this error there is a good chance that your
# database lock file is owned by root
$ oozie job -oozie http://localhost:11000/oozie -config \
examples/apps/sqoop/job.properties -run
Error: E0607 : E0607: Other error in operation [<openjpa-1.2.1-r752877:753278
fatal store error> org.apache.openjpa.persistence.RollbackException:
The transaction has been rolled back. See the nested exceptions for
details on the errors that occurred.], {1}
# fix this as follows
$ sudo chown oozie:oozie /var/lib/oozie/oozie-db/db.lck
# and restart the oozie server
$ sudo su - oozie -s /usr/lib/oozie/bin/oozie-stop.sh
$ sudo su - oozie -s /usr/lib/oozie/bin/oozie-start.sh
|
HBase là một kho lưu trữ key-value hiệu năng cao. Nếu trường hợp sử dụng của bạn cần có khả năng mở rộng quy mô và chỉ cần kiểu giao dịch cơ sở dữ liệu tương đương như các giao dịch giao kết tự động, thì rất có thể HBase là công nghệ cần dùng. HBase không phải là một cơ sở dữ liệu. Tên của nó không thích hợp vì đối với một số người, thuật ngữ base ngụ ý là cơ sở dữ liệu. Nó thực sự làm việc tuyệt vời cho các kho lưu trữ key-value hiệu năng cao. Có một số sự chồng chéo giữa chức năng của HBase, Informix, DB2 và các cơ sở dữ liệu quan hệ khác. Đối với các giao dịch ACID, tuân thủ SQL đầy đủ, và nhiều chỉ mục thì một cơ sở dữ liệu quan hệ truyền thống sẽ là sự lựa chọn hiển nhiên.
Bài tập viết mã cuối cùng này nhằm đem đến sự hiểu biết cơ bản về HBase. Bài tập này được thiết kế đơn giản và hoàn toàn không đại diện cho phạm vi chức năng của HBase. Hãy sử dụng ví dụ này để hiểu một số trong những khả năng cơ bản trong HBase. "HBase, The Definitive Guide" (HBase một hướng dẫn đáng tin cậy) của Lars George, là cuốn sách bắt buộc phải đọc nếu bạn có kế hoạch thực hiện hoặc loại bỏ HBase trong trường hợp sử dụng cụ thể của mình.
Ví dụ cuối cùng, được hiển thị trong Liệt kê 23 và 24, sử dụng giao diện REST được cung cấp với HBase để chèn các cặp khóa-các giá trị vào một bảng HBase. Chạy bài thử nghiệm bằng curl.
Liệt kê 23. Tạo một bảng HBase và chèn vào một hàng
# enter the command line shell for hbase
$ hbase shell
HBase Shell; enter 'help<RETURN> for list of supported commands.
Type "exit<RETURN> to leave the HBase Shell
Version 0.90.6-cdh3u4, r, Mon May 7 13:14:00 PDT 2012
# create a table with a single column family
hbase(main):001:0> create 'mytable', 'mycolfamily'
# if you get errors from hbase you need to fix the
# network config
# here is a sample of the error:
ERROR: org.apache.hadoop.hbase.ZooKeeperConnectionException: HBase
is able to connect to ZooKeeper but the connection closes immediately.
This could be a sign that the server has too many connections
(30 is the default). Consider inspecting your ZK server logs for
that error and then make sure you are reusing HBaseConfiguration
as often as you can. See HTable's javadoc for more information.
# fix networking:
# add the eth0 interface to /etc/hosts with a hostname
$ sudo su -
# ifconfig | grep addr
eth0 Link encap:Ethernet HWaddr 00:0C:29:8C:C7:70
inet addr:192.168.1.134 Bcast:192.168.1.255 Mask:255.255.255.0
Interrupt:177 Base address:0x1400
inet addr:127.0.0.1 Mask:255.0.0.0
[root@myhost ~]# hostname myhost
[root@myhost ~]# echo "192.168.1.134 myhost" >gt; /etc/hosts
[root@myhost ~]# cd /etc/init.d
# now that the host and address are defined restart Hadoop
[root@myhost init.d]# for i in hadoop*
> do
> ./$i restart
> done
# now try table create again:
$ hbase shell
HBase Shell; enter 'help<RETURN> for list of supported commands.
Type "exit<RETURN> to leave the HBase Shell
Version 0.90.6-cdh3u4, r, Mon May 7 13:14:00 PDT 2012
hbase(main):001:0> create 'mytable' , 'mycolfamily'
0 row(s) in 1.0920 seconds
hbase(main):002:0>
# insert a row into the table you created
# use some simple telephone call log data
# Notice that mycolfamily can have multiple cells
# this is very troubling for DBAs at first, but
# you do get used to it
hbase(main):001:0> put 'mytable', 'key123', 'mycolfamily:number','6175551212'
0 row(s) in 0.5180 seconds
hbase(main):002:0> put 'mytable', 'key123', 'mycolfamily:duration','25'
# now describe and then scan the table
hbase(main):005:0> describe 'mytable'
DESCRIPTION ENABLED
{NAME => 'mytable', FAMILIES => [{NAME => 'mycolfam true
ily', BLOOMFILTER => 'NONE', REPLICATION_SCOPE => '
0', COMPRESSION => 'NONE', VERSIONS => '3', TTL =>
'2147483647', BLOCKSIZE => '65536', IN_MEMORY => 'f
alse', BLOCKCACHE => 'true'}]}
1 row(s) in 0.2250 seconds
# notice that timestamps are included
hbase(main):007:0> scan 'mytable'
ROW COLUMN+CELL
key123 column=mycolfamily:duration,
timestamp=1346868499125, value=25
key123 column=mycolfamily:number,
timestamp=1346868540850, value=6175551212
1 row(s) in 0.0250 seconds
|
Liệt kê 24. Sử dụng giao diện REST của Hbase
# HBase includes a REST server
$ hbase rest start -p 9393 &
# you get a bunch of messages....
# get the status of the HBase server
$ curl http://localhost:9393/status/cluster
# lots of output...
# many lines deleted...
mytable,,1346866763530.a00f443084f21c0eea4a075bbfdfc292.
stores=1
storefiless=0
storefileSizeMB=0
memstoreSizeMB=0
storefileIndexSizeMB=0
# now scan the contents of mytable
$ curl http://localhost:9393/mytable/*
# lines deleted
12/09/05 15:08:49 DEBUG client.HTable$ClientScanner:
Finished with scanning at REGION =>
# lines deleted
<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
<CellSet><Row key="a2V5MTIz">
<Cell timestamp="1346868499125" column="bXljb2xmYW1pbHk6ZHVyYXRpb24=">MjU=</Cell>
<Cell timestamp="1346868540850" column="bXljb2xmYW1pbHk6bnVtYmVy">NjE3NTU1MTIxMg==</Cell>
<Cell timestamp="1346868425844" column="bXljb2xmYW1pbHk6bnVtYmVy">NjE3NTU1MTIxMg==</Cell>
</Row></CellSet>
# the values from the REST interface are base64 encoded
$ echo a2V5MTIz | base64 -d
key123
$ echo bXljb2xmYW1pbHk6bnVtYmVy | base64 -d
mycolfamily:number
# The table scan above gives the schema needed to insert into the HBase table
$ echo RESTinsertedKey | base64
UkVTVGluc2VydGVkS2V5Cg==
$ echo 7815551212 | base64
NzgxNTU1MTIxMgo=
# add a table entry with a key value of "RESTinsertedKey" and
# a phone number of "7815551212"
# note - curl is all on one line
$ curl -H "Content-Type: text/xml" -d '<CellSet>
<Row key="UkVTVGluc2VydGVkS2V5Cg==">
<Cell column="bXljb2xmYW1pbHk6bnVtYmVy">NzgxNTU1MTIxMgo=<Cell>
<Row><CellSet> http://192.168.1.134:9393/mytable/dummykey
12/09/05 15:52:34 DEBUG rest.RowResource: POST http://192.168.1.134:9393/mytable/dummykey
12/09/05 15:52:34 DEBUG rest.RowResource: PUT row=RESTinsertedKey\x0A,
families={(family=mycolfamily,
keyvalues=(RESTinsertedKey\x0A/mycolfamily:number/9223372036854775807/Put/vlen=11)}
# trust, but verify
hbase(main):002:0> scan 'mytable'
ROW COLUMN+CELL
RESTinsertedKey\x0A column=mycolfamily:number,timestamp=1346874754883,value=7815551212\x0A
key123 column=mycolfamily:duration, timestamp=1346868499125, value=25
key123 column=mycolfamily:number, timestamp=1346868540850, value=6175551212
2 row(s) in 0.5610 seconds
# notice the \x0A at the end of the key and value
# this is the newline generated by the "echo" command
# lets fix that
$ printf 8885551212 | base64
ODg4NTU1MTIxMg==
$ printf mykey | base64
bXlrZXk=
# note - curl statement is all on one line!
curl -H "Content-Type: text/xml" -d '<CellSet><Row key="bXlrZXk=">
<Cell column="bXljb2xmYW1pbHk6bnVtYmVy">ODg4NTU1MTIxMg==<Cell>
<Row><CellSet>
http://192.168.1.134:9393/mytable/dummykey
# trust but verify
hbase(main):001:0> scan 'mytable'
ROW COLUMN+CELL
RESTinsertedKey\x0A column=mycolfamily:number,timestamp=1346875811168,value=7815551212\x0A
key123 column=mycolfamily:duration, timestamp=1346868499125, value=25
key123 column=mycolfamily:number, timestamp=1346868540850, value=6175551212
mykey column=mycolfamily:number, timestamp=1346877875638, value=8885551212
3 row(s) in 0.6100 seconds
|
Sử dụng Pig: Nối dữ liệu Informix và dữ liệu DB2
Pig là một ngôn ngữ thủ tục. Cũng giống như Hive, nó tạo mã MapReduce ngầm bên dưới vỏ bọc. Tính dễ sử dụng của Hadoop sẽ còn tiếp tục được cải thiện khi nhiều dự án trở nên có sẵn. Cũng giống như một số người trong chúng ta thực sự thích dòng lệnh, có một số giao diện người dùng đồ họa làm việc rất tốt với Hadoop.
Liệt kê 19 cho thấy mã Pig được sử dụng để nối bảng customer và bảng staff từ ví dụ trước.
Liệt kê 19. Ví dụ Pig để nối bảng Informix với bảng DB2
$ pig
grunt> staffdb2 = load 'staff' using PigStorage(',')
>> as ( id, name, dept, job, years, salary, comm );
grunt> custifx2 = load 'customer' using PigStorage(',') as
>> (cn, fname, lname, company, addr1, addr2, city, state, zip, phone)
>> ;
grunt> joined = join custifx2 by cn, staffdb2 by id;
# to make pig generate a result set use the dump command
# no work has happened up till now
grunt> dump joined;
2012-08-11 21:24:51,848 [main] INFO org.apache.pig.tools.pigstats.ScriptState
- Pig features used in the script: HASH_JOIN
2012-08-11 21:24:51,848 [main] INFO org.apache.pig.backend.hadoop.executionengine
.HExecutionEngine - pig.usenewlogicalplan is set to true.
New logical plan will be used.
HadoopVersion PigVersion UserId StartedAt FinishedAt Features
0.20.2-cdh3u4 0.8.1-cdh3u4 cloudera 2012-08-11 21:24:51
2012-08-11 21:25:19 HASH_JOIN
Success!
Job Stats (time in seconds):
JobId Maps Reduces MaxMapTime MinMapTIme AvgMapTime
MaxReduceTime MinReduceTime AvgReduceTime Alias Feature Outputs
job_201208111415_0006 2 1 8 8 8 10 10 10
custifx,joined,staffdb2 HASH_JOIN hdfs://0.0.0.0/tmp/temp1785920264/tmp-388629360,
Input(s):
Successfully read 35 records from: "hdfs://0.0.0.0/user/cloudera/staff"
Successfully read 28 records from: "hdfs://0.0.0.0/user/cloudera/customer"
Output(s):
Successfully stored 2 records (377 bytes) in:
"hdfs://0.0.0.0/tmp/temp1785920264/tmp-388629360"
Counters:
Total records written : 2
Total bytes written : 377
Spillable Memory Manager spill count : 0
Total bags proactively spilled: 0
Total records proactively spilled: 0
Job DAG:
job_201208111415_0006
2012-08-11 21:25:19,145 [main] INFO
org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.MapReduceLauncher - Success!
2012-08-11 21:25:19,149 [main] INFO org.apache.hadoop.mapreduce.lib.
input.FileInputFormat - Total input paths to process : 1
2012-08-11 21:25:19,149 [main] INFO org.apache.pig.backend.hadoop.
executionengine.util.MapRedUtil - Total input paths to process : 1
(110,Roy ,Jaeger ,AA Athletics ,520 Topaz Way
,null,Redwood City ,CA,94062,415-743-3611 ,110,Ngan,15,Clerk,5,42508.20,206.60)
(120,Fred ,Jewell ,Century Pro Shop ,6627 N. 17th Way
,null,Phoenix ,AZ,85016,602-265-8754
,120,Naughton,38,Clerk,null,42954.75,180.00)
|
Sử dụng Hive: Nối dữ liệu Informix và dữ liệu DB2
Có một trường hợp sử dụng thú vị cần nối dữ liệu từ Informix đến DB2. Không hứng thú lắm đối với hai bảng tầm thường, nhưng sẽ là một thắng lợi vĩ đại với nhiều terabyte hoặc petabyte dữ liệu.
Có hai cách tiếp cận cơ bản để nối các nguồn dữ liệu khác nhau. Cứ giữ nguyên dữ liệu ở đâu ở đó và sử dụng công nghệ liên hợp dữ liệu đối lập với việc di chuyển dữ liệu đến một kho lưu trữ duy nhất để thực hiện phép nối. Yếu tố kinh tế và hiệu năng của Hadoop làm cho việc di chuyển dữ liệu vào HDFS và thực hiện công việc nặng nhọc với MapReduce trở thành một sự lựa chọn dễ dàng. Các giới hạn băng thông mạng tạo ra một rào cản chính nếu cố gắng nối dữ liệu ở nguyên chỗ cũ bằng một công nghệ kiểu liên hợp. Để biết thêm thông tin về liên hợp dữ liệu, hãy xem phần Tài nguyên.
Hive cung cấp một tập hợp con của SQL để hoạt động trên một cụm. Nó không cung cấp ngữ nghĩa giao dịch. Nó không phải là một sự thay thế cho Informix hoặc DB2. Nếu bạn có một công việc nặng nề nào đó dưới dạng các phép nối bảng, ngay cả khi bạn có một số bảng nhỏ hơn nhưng cần phải thực hiện các tích Đề-các (Cartesian) khó chịu, thì Hadoop là công cụ nên dùng.
Để sử dụng ngôn ngữ truy vấn Hive, cần có một tập hợp con của SQL được gọi là siêu dữ liệu bảng Hiveql. Bạn có thể định nghĩa siêu dữ liệu này dựa vào các tệp hiện có trong HDFS. Sqoop cung cấp một lối tắt thuận tiện với tùy chọn create-hive-table (tạo bảng hive).
Những người dùng MySQL sẽ thấy thoải mái khi sửa lại các ví dụ được hiển thị trong Liệt kê 18 cho phù hợp. Một bài tập thú vị là nối MySQL hoặc bất kỳ các bảng cơ sở dữ liệu quan hệ khác nào, với các bảng tính lớn.
Liệt kê 18. Nối bảng informix.customer với bảng db2.staff
# import the customer table into Hive
$ sqoop import --driver com.informix.jdbc.IfxDriver \
--connect \
"jdbc:informix-sqli://myhost:54321/stores_demo:informixserver=ifx;user=me;password=you" \
--table customer
# now tell hive where to find the informix data
# to get to the hive command prompt just type in hive
$ hive
Hive history file=/tmp/cloudera/yada_yada_log123.txt
hive>
# here is the hiveql you need to create the tables
# using a file is easier than typing
create external table customer (
cn int,
fname string,
lname string,
company string,
addr1 string,
addr2 string,
city string,
state string,
zip string,
phone string)
ROW FORMAT DELIMITED FIELDS TERMINATED BY ','
LOCATION '/user/cloudera/customer'
;
# we already imported the db2 staff table above
# now tell hive where to find the db2 data
create external table staff (
id int,
name string,
dept string,
job string,
years string,
salary float,
comm float)
ROW FORMAT DELIMITED FIELDS TERMINATED BY ','
LOCATION '/user/cloudera/staff'
;
# you can put the commands in a file
# and execute them as follows:
$ hive -f hivestaff
Hive history file=/tmp/cloudera/hive_job_log_cloudera_201208101502_2140728119.txt
OK
Time taken: 3.247 seconds
OK
10 Sanders 20 Mgr 7 98357.5 NULL
20 Pernal 20 Sales 8 78171.25 612.45
30 Marenghi 38 Mgr 5 77506.75 NULL
40 O'Brien 38 Sales 6 78006.0 846.55
50 Hanes 15 Mgr 10 80
... lines deleted
# now for the join we've all been waiting for :-)
# this is a simple case, Hadoop can scale well into the petabyte range!
$ hive
Hive history file=/tmp/cloudera/hive_job_log_cloudera_201208101548_497937669.txt
hive> select customer.cn, staff.name,
> customer.addr1, customer.city, customer.phone
> from staff join customer
> on ( staff.id = customer.cn );
Total MapReduce jobs = 1
Launching Job 1 out of 1
Number of reduce tasks not specified. Estimated from input data size: 1
In order to change the average load for a reducer (in bytes):
set hive.exec.reducers.bytes.per.reducer=number
In order to limit the maximum number of reducers:
set hive.exec.reducers.max=number
In order to set a constant number of reducers:
set mapred.reduce.tasks=number
Starting Job = job_201208101425_0005,
Tracking URL = http://0.0.0.0:50030/jobdetails.jsp?jobid=job_201208101425_0005
Kill Command = /usr/lib/hadoop/bin/hadoop
job -Dmapred.job.tracker=0.0.0.0:8021 -kill job_201208101425_0005
2012-08-10 15:49:07,538 Stage-1 map = 0%, reduce = 0%
2012-08-10 15:49:11,569 Stage-1 map = 50%, reduce = 0%
2012-08-10 15:49:12,574 Stage-1 map = 100%, reduce = 0%
2012-08-10 15:49:19,686 Stage-1 map = 100%, reduce = 33%
2012-08-10 15:49:20,692 Stage-1 map = 100%, reduce = 100%
Ended Job = job_201208101425_0005
OK
110 Ngan 520 Topaz Way Redwood City 415-743-3611
120 Naughton 6627 N. 17th Way Phoenix 602-265-8754
Time taken: 22.764 seconds
|
Sẽ đẹp hơn nhiều khi bạn sử dụng Hue để có một giao diện đồ họa trong trình duyệt, như thể hiện trong Hình 9, 10 và 11.
Hình 9. Giao diện người dùng đồ họa (GUI) Beeswax của Hue cho Hive trong CDH4, xem truy vấn Hiveql
Hình 10. Giao diện người dùng đồ họa (GUI) Beeswax của Hue cho Hive, xem truy vấn Hiveql
Hình 11. Giao diện người dùng đồ họa (GUI) Beeswax của Hue, xem kết quả của phép nối Informix-DB2
Đăng ký:
Bài đăng (Atom)